You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Phase A: Stabilize float32 AMICATorchNG via Kahan + mixed precision #75
Enabler phase of the Apple-Silicon GPU epic. Apple GPUs have no FP64, so all Apple-GPU
acceleration requires a stable float32 AMICA. #70 established that float32 diverges to
NaN on the full 30504-sample data (precision-driven instability, not a divide-by-zero)
and that computing the responsibilities in float64 alone does not fix it.
Approach
Kahan / compensated summation for the sufficient-stat accumulation (~30504 samples
per block/component - the prime float32 precision sink Stabilize float32 AMICATorchNG for the GPU fast path #70 never touched). Near-float64
accuracy with an error bound independent of the term count N, at ~2-3x the sum cost.
Mixed precision: float64 for the density |y|^rho, the log-likelihood, and the
stat accumulation; float32 for the matmuls. Guard behind dtype so the default float64
path is untouched.
Acceptance
float32 (or mixed-precision) fit converges on the full sample data (32ch x 30504) across
=5 seeds, Newton on and off, with no NaN.
Final LL within parity tolerance of the float64 fit.
Tests use real sample data (seed sweep) under tests/torch_tests/.
References
Supersedes #70 (parked). Research: .context/mps_pathways.md Pathway A. Perf-sensitive
regions identified in .context/issue-63/perf_findings.md sections 1 and 4.
Done via PR #78 (squash-merged into the epic branch feature/issue-74-epic-apple-gpu, commit 6d8a002). Root cause was a per-element float32 0/0 in the mu denominator ufp/y (a sample rounding an activation to exactly 0), not summation precision; a one-line divide-by-zero guard fixes it, bit-identical in float64 and needing no float64 (so it holds on MPS). Verified across 5 seeds x Newton on/off, LL matches float64 to ~5 sig digits. Phase A of epic #74.
Context
Enabler phase of the Apple-Silicon GPU epic. Apple GPUs have no FP64, so all Apple-GPU
acceleration requires a stable float32 AMICA. #70 established that float32 diverges to
NaN on the full 30504-sample data (precision-driven instability, not a divide-by-zero)
and that computing the responsibilities in float64 alone does not fix it.
Approach
per block/component - the prime float32 precision sink Stabilize float32 AMICATorchNG for the GPU fast path #70 never touched). Near-float64
accuracy with an error bound independent of the term count N, at ~2-3x the sum cost.
|y|^rho, the log-likelihood, and thestat accumulation; float32 for the matmuls. Guard behind dtype so the default float64
path is untouched.
Acceptance
tests/torch_tests/.References
Supersedes #70 (parked). Research:
.context/mps_pathways.mdPathway A. Perf-sensitiveregions identified in
.context/issue-63/perf_findings.mdsections 1 and 4.