Context
Higher-ceiling Apple-GPU path of the epic. MLX (Apple's array framework) beats PyTorch
MPS by ~2-3x on the same hardware, is Apple-native with a real compiler / lazy graph, and
can fuse the per-block work that PyTorch MPS runs eager and largely unfused. GPU arrays are
float32-only, so this depends on Phase A.
Approach
- Port the
AMICATorchNG per-block E/M-step to MLX arrays behind a backend switch, keeping
the PyTorch path as the float64 parity / CUDA backend.
- Reuse the Phase A mixed-precision strategy (float64 accumulation on CPU where MLX allows,
float32 on GPU).
- Keep MLX an optional dependency (not required for the core install).
Acceptance
- MLX backend produces results equivalent to the PyTorch float32 backend on the sample data
(matched LL within tolerance).
- Behavior-validated (no bit-exact float64 oracle is possible on an Apple GPU).
- Optional-dependency install path documented.
Dependencies
Blocked by Phase A (MLX GPU is float32-only). Research: .context/mps_pathways.md
Pathway C.
Context
Higher-ceiling Apple-GPU path of the epic. MLX (Apple's array framework) beats PyTorch
MPS by ~2-3x on the same hardware, is Apple-native with a real compiler / lazy graph, and
can fuse the per-block work that PyTorch MPS runs eager and largely unfused. GPU arrays are
float32-only, so this depends on Phase A.
Approach
AMICATorchNGper-block E/M-step to MLX arrays behind a backend switch, keepingthe PyTorch path as the float64 parity / CUDA backend.
float32 on GPU).
Acceptance
(matched LL within tolerance).
Dependencies
Blocked by Phase A (MLX GPU is float32-only). Research:
.context/mps_pathways.mdPathway C.