Skip to content

Phase A: Stabilize float32 AMICATorchNG via ufp/y guard - #78

Merged
neuromechanist merged 2 commits into
feature/issue-74-epic-apple-gpufrom
feature/issue-75-phaseA-float32-stabilize
Jul 8, 2026
Merged

neuromechanist merged 2 commits into
feature/issue-74-epic-apple-gpufrom
feature/issue-75-phaseA-float32-stabilize

Conversation

@neuromechanist

Copy link
Copy Markdown
Member

Summary

float32 AMICATorchNG diverged to NaN on the full 30504-sample sample EEG across every seed (Newton on and off, crashing iter ~9-105), while float64 converged. This blocked the whole Apple-GPU roadmap (epic #74): Apple GPUs have no FP64, so every MPS/MLX pathway is gated on a stable float32 AMICA.

Root cause (a per-element divide-by-zero, not summation precision). The mu denominator is sbeta*sum(ufp/y) with ufp = u*fp. At a sample sitting on a mixture mean, float32 rounds the scaled activation y to exactly 0, and the score fp(0)=0 for every family, so that term is 0/0 = NaN and one NaN summand poisons the whole dmu_d. float64 never lands y on exact 0, hence float64-only.

Fix: a one-line guard, ufp / where(y==0, 1, y) (contributing 0 for the measure-zero sample). It is a no-op in float64 (bit-identical, so single-model #24 Fortran parity is preserved) and needs no float64, so it also stabilizes the MPS/float32 path.

Diagnosis (why not compensated summation)

The plan led with an inter-block summation-precision hypothesis; diagnostics on the real data refuted it and landed on the pre-registered fallback branch:

  • Accumulating the block sufficient statistics in float64, and Neumaier compensated summation, did not help (both still diverged, at the same iterations).
  • Computing the density |y|^rho/log_pdf in float64 did not help (nor, per Stabilize float32 AMICATorchNG for the GPU fast path #70, the responsibilities).
  • Only guarding the ufp/y division does. So the earlier "needs mixed precision, payoff ~1.5-2x" conclusion is superseded; the fix needs no float64 at all.

Test plan

  • New tests/torch_tests/test_ng_float32_stability.py: full-data float32 fit converges across 5 seeds spanning Newton on/off (finite LL, non-degenerate stop), and float32 LL matches the in-test float64 fit within 0.05 (observed: ~5 significant digits). Plus a fast fp(0)=0 invariant test and an MPS smoke test (self-skips in CI).
  • Full non-slow torch suite green: 124 passed, 5 skipped, including the NumPy sufficient-stats parity (1e-8) and byte-identity anchors -> float64 confirmed bit-identical.
  • ruff check + ruff format --check clean.

Docs updated

.context/issue-63/perf_findings.md (section 4 corrected: it IS a divide-by-zero, per-element), .context/mps_pathways.md (Pathway A marked DONE, mechanism corrected), AGENTS.md, benchmarks/benchmark_gpu.py.

Closes #75
Part of epic #74

float32 diverged to NaN on the full 30504-sample data across every seed
(Newton on and off), while float64 converged. Root cause: the mu denominator
sbeta*sum(ufp/y) (ufp=u*fp); at a sample sitting on a mixture mean, float32
rounds the scaled activation y to exactly 0, and fp(0)=0 for every family, so
that term is 0/0=NaN and one NaN summand poisons dmu_d. float64 never rounds y
to exactly 0.

Diagnostics ruled out summation precision (accumulating block partials in
float64, and Neumaier compensated summation, did not help) and the density /
responsibilities (float64 there did not help either). Only guarding the ufp/y
division does. The guard (ufp / where(y==0, 1, y), contributing 0 for the
measure-zero sample) is a no-op in float64 (y is never exactly 0), so
single-model #24 parity stays bit-identical, and it needs no float64, so it
also stabilizes the MPS/float32 path (Apple GPUs have no FP64) -- epic #74
Phase A.

Tested: float32 now converges across 5 seeds x Newton on/off on the real
sample EEG, matching the float64 LL to ~5 significant digits; full non-slow
torch suite green (124 passed), including the NumPy-parity and byte-identity
anchors. Docs updated (perf_findings, mps_pathways, AGENTS, benchmark_gpu).
- Correct the guard comment: the true ufp/y limit is NOT 0 (nonzero constant at
  rho=2, integrable singularity diverging for rho<2), so the guard drops an
  unrepresentable singular term, not a removable zero. Measured: it fires <=1
  sample/iteration on the sample EEG (5 of 150 iters), and float32 still matches
  the float64 LL to ~5 sig digits -- a bounded, negligible bias.
- Scope the "fp(0)=0 for every family" claim to the supported rho>=1 (for rho<1
  the GG fp is itself NaN at 0; out of scope, default minrho=1.0).
- test: reference AMICATorchNG._DEGENERATE_STOP_REASONS instead of duplicating
  the ("nan_ll","singular_ll") tuple; shorten the float64-tracking quality check
  to 100 iters (it tracks float64 from iter 1; the 150-iter sweep remains the
  regression guard).
- docs: finish the mps_pathways.md update -- intro to past tense, and Pathways B
  and C no longer gate on Pathway A as unmet (it is done, #75).
@neuromechanist

Copy link
Copy Markdown
Member Author

Review summary (4 Sonnet reviewers: code, silent-failure, tests, comments/docs)

Mutation-tested and independently verified: removing the guard fails all 5 test_float32_stable_on_full_data cases (crash iter 26-57, inside the 150 budget); with it, the full non-slow torch suite is green (124 passed) including the NumPy sufficient-stats parity (1e-8) and byte-identity anchors -> float64 confirmed bit-identical.

Addressed (commit ad90407)

  • Guard comment was mathematically overstated (silent-failure + comment + code reviewers). It implied 0 is the correct value because ufp==0 at y==0. Corrected: the true term ufp/y = u*rho*|y|^(rho-2) is not 0 in the limit -- a nonzero constant at rho=2, and an integrable singularity diverging as y->0 for rho<2. Once float32 rounds y to exactly 0 the real contribution is unrepresentable; substituting 0 drops that one sample rather than poisoning dmu_d with a NaN. Measured firing rate: it fires <=1 sample per iteration on the sample EEG (5 of 150 iterations; peak 3.4e-7 of all y), and float32 still matches the float64 LL to ~5 significant digits -- a bounded, negligible bias.
  • fp(0)=0 for every family scoped to the supported rho>=1 (comment + code reviewers): for an unsupported rho<1 the GG fp is itself NaN at 0, orthogonal to this guard (default minrho=1.0).
  • mps_pathways.md leftover contradictions (comment reviewer): intro moved to past tense, and Pathways B/C no longer gate on Pathway A as unmet (it is done).
  • Test hygiene (test reviewer): reference AMICATorchNG._DEGENERATE_STOP_REASONS instead of duplicating the tuple; the float64-tracking quality check shortened to 100 iters (it tracks float64 from iter 1), keeping the 150-iter sweep as the regression guard.

Considered and not changed (with rationale)

  • Per-fire logging for the guard (silent-failure, HIGH). Declined: unlike the sibling guards (dead-model, NaN-rho reset), which log genuine trouble, this fires for a benign measure-zero event (a sample landing exactly on a mixture mean) at ~0.03x/iteration; a per-iteration warning would be alarm-fatigue noise. The behavior is now documented in-code and bounded by the test_float32_ll_matches_float64 acceptance test, which satisfies the no-hidden-failures intent.
  • Full 5x2 seed/Newton cross-product (test reviewer confirmed sound): the 0/0 is in the always-on E-step accumulation, Newton-independent (mutation failed identically for both), so 5 distinct seeds spanning both settings is adequate coverage at lower CI cost.
  • CI runtime (test reviewer, 6/10 "watch, don't mark slow"): marking these slow would disable the only automated guard against this regression, so the sweep stays in the gate; the quality-check trim above is the safe saving.

Process note

During the parallel review, one write-capable reviewer left a stray uncommitted safe_y = y mutation in the shared worktree (regression-detection test) that briefly reverted the guard. HEAD (48f1ff2) was always correct; the working tree was restored to match before committing. Next phase I will run write-capable reviewers with worktree isolation.

@neuromechanist
neuromechanist merged commit 6d8a002 into feature/issue-74-epic-apple-gpu Jul 8, 2026
5 checks passed
@neuromechanist
neuromechanist deleted the feature/issue-75-phaseA-float32-stabilize branch July 8, 2026 08:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant