Skip to content

Phase 15: re-measure parity under the finished epic (#351) - #356

Merged
neuromechanist merged 64 commits into
feature/issue-324-epic-mlx-first-classfrom
feature/issue-351-phase15-final-parity
Sep 24, 2026
Merged

neuromechanist merged 64 commits into
feature/issue-324-epic-mlx-first-classfrom
feature/issue-351-phase15-final-parity

Conversation

@neuromechanist

@neuromechanist neuromechanist commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Summary

Re-measures every user-facing parity figure with the finished epic #324 code (09b8247 + #353; merged with Phase 17, bb22df6, before review), against the pinned v0.3.3 native binary, and updates the documents that quote them: paper.md (Table 1, Validation, figure), docs/guides/validation.md, docs/guides/amica-differences.md, docs/guides/backends.md, ADR 0003 (new dated section), AGENTS.md, .context/progress_summary.md, benchmarks/README_dimsweep.md, the changelog, and the regenerated docs/assets/figures/multimodel-ensemble.png. Scripts and raw outputs are under .context/issue-351/; the findings record (findings.md) is committed by the lead.

No library code changes. benchmarks/reproduce_table1.py now pins every reference setting explicitly (REFERENCE_SETTINGS), and its ensemble's lrate, so Phase 17's default changes cannot move the Table 1 protocol. Before Phase 17 the pinned and unpinned calls wrote byte-identical input.param files at all 8 call sites (.context/issue-351/pin_check.py). After the Phase 17 merge each pinned call writes the same settings as before, in Phase 17's key order and with do_approx_sphere 1 added, the binary's compiled value (now pinned in REFERENCE_SETTINGS as well); the binary gives byte-identical output files for both forms of the file (pin_run_check.py, bundled sample, three configurations). Without the pins, 15 settings of the native calls would have moved to Phase 17's defaults.

Part of #324, closes #351 on the epic merge.

Old vs new

Figure Before the epic Now Record
Harness (100 it, independent starts): dLL torch / NumPy / MLX 2.9e-5 / 2.9e-5 / 3.2e-5 2.71e-4 / 2.72e-4 / 2.76e-4 (reference's own seed-to-seed LL sd 2.6e-4) raw/harness_run1, raw/harness_gap_head.json
Harness: mean / min corr, Amari 0.9992 / 0.9935, 0.0037 0.9991 / 0.9918, 0.0038 raw/harness_run1
Harness from a shared start: dLL, corr, Amari 2.35e-4, 0.99999, 5.8e-4 (e38aa11) 1.6e-6, 0.99999993, 3.9e-5 raw/harness_gap_*.json
Harness runtimes Fortran / torch / NumPy / MLX (s) 10.1 / 17.9 / 32.5 / 3.2 (same session) 10.0 / 18.1 / 32.6 / 3.0 raw/harness_*
200-it amicaout fixture: dLL, corr, Amari 2.2-2.3e-4, 0.9973, 6.3e-3 (pre Phase 11) 1.3-1.4e-4, 0.9983, 4.8e-3 raw/fixture_parity_head.txt
MLX f32 vs torch f64, same start, 100 it not characterized dLL 5.0e-6, corr 0.99999991, Amari 5.0e-5 raw/precision_agreement.json
Table 1 Amari (bundled, 5 pairs) 0.006 0.011 (one reference run in a lower-LL basin); 0.004 over 50 pairs (reference vs itself 0.005) raw/table1_bundled, raw/bundled_basins
Score / sufficient statistics ~1e-15 8.9e-16 / 1.8e-15; 4.5e-13 abs (3e-16 relative) raw/table1_bundled
Multi corr cross; within-Fortran 0.65; 0.64 0.632; 0.626 (perm p 0.88) raw/table1_bundled, multimodel_summary.json
Multi Amari cross; within-Fortran 0.163; 0.174 (p > 0.999) 0.172; 0.166 (p 0.051) same
Multi ensemble LL Fortran; pamica -3.3539; -3.3629 (KS p 6e-5) -3.3543; -3.3541 (KS p 0.83) same, raw/multimodel (pre-epic refit -3.3627)
keep_best: sd ratio at 100 it, restores 12.7x -> 2.0x 1.0x with or without; 1 restore in 20 fits (300 it only, gain 4.2e-6) raw/keep_best
Mean LL gap at 100 / 200 / 300 it -0.009 / 0 / +0.002 +8.1e-4 / +2.1e-4 / +1.0e-4 raw/keep_best
Cross-backend 25-it LL (dimsweep) max pairwise ~0.003 1e-5 at 32/48 ch; up to 1.1e-3 at 70 ch (k 6, round-off sensitivity 4.4e-4) raw/dimsweep
Newton, independent seeds (70 ch) seed 42: 10-11/70 below 0.9; others ~0.995 seed 42: 8/70 vs F1; others 0.995-0.996; reference vs itself 0.985 (2/70) raw/newton
Newton, same start 0.9974 (min 0.9475) 0.9999998 (min 0.9999964), dLL 7.6e-8 raw/newton
share_comps e2e (300 it, thr 0.95) 3 merges, LL -3.3416 1 merge ( cos
Early scans at it 8 / 20 (0.95, 0.99) 32, 24 / 24, 5 30, 17 / 22, 5 raw/share_comps
Reference-formula scan on its own state 32, 32, 30 32, 32, 29 raw/share_comps
Table 1 external LL (on002718, 5 seeds) gap ~0.0003 at -3.6993 Fortran -3.699346, pamica -3.699341, gap 5.6e-6 raw/table1_external
Table 1 external correlation 0.998 (Fortran vs Fortran ~0.999) 0.9996, min component 0.960 (Fortran vs Fortran 0.998); Amari 0.0013 (0.0025) raw/table1_external

Notes

  • The harness dLL and the 70-channel dimsweep spread are numerically larger than before; both were reported to the lead with evidence (start-to-start spread and round-off amplification) before the docs were edited.
  • The old ensemble LL gap came partly from the pre-Apply the LL-decrease response before the update #339 min_dll check (7 of 20 pre-epic fits stopped early) and partly from that code's update rule; the "convergence speed" explanation is removed.
  • The AMICATorchNG diverges from Fortran on weak components at 2000 iters, despite Newton making the optimum sharply unique #145 reference runs came from the gfortran amica15_linux build, which draws a new start on each unseeded run (checked on hallu), so both reference-vs-reference figures are reported side by side.
  • Pre-epic diagnostics use e38aa11 (Phase 6), loaded as a separate tree; the orientation of the seeded state was checked by the first-iteration LL (9e-16).

Test plan

  • uv run ruff check ., uv run ruff format --check ., uv run ty check ., typos
  • uv run pytest -n 8 --no-cov with CI=1: 1955 passed, 40 skipped, 3 xfailed before the Phase 17 merge; 1974 passed, 40 skipped, 3 xfailed after it
  • AMICA_RUN_FORTRAN=1 on test_fortran_param_forwarding.py and the two early-merge oracles: 37 passed, before and after the Phase 17 merge
  • uv run --extra docs mkdocs build --strict
  • pin_check.py: byte-identical input.param at all 8 native call sites before Phase 17, with a negative control; after it, the same settings at all 8 (--against the saved files), plus do_approx_sphere 1
  • pin_run_check.py: the binary's outputs are byte-identical for the pre- and post-Phase-17 forms of the pinned files (wall-clock times removed from out.txt)
  • External tier numbers integrated (lead's run at 09b8247, 5/5 seeds, 26185 s)
  • Epic merged, including Phase 16 (3fbe1ab) and Phase 17 (bb22df6); one conflict in the Newton section of validation.md, resolved with Phase 17's wording of the do_newton defaults
  • Review findings (numbers, docs claims, paper) addressed; Figure 1 panels now use line styles as well as color

neuromechanist and others added 30 commits September 23, 2026 12:58
…-class' into feature/issue-351-phase15-final-parity

# Conflicts:
#	AGENTS.md
…parity' into feature/issue-351-phase15-final-parity
@neuromechanist
neuromechanist marked this pull request as ready for review September 24, 2026 05:43
@neuromechanist
neuromechanist merged commit 880d71d into feature/issue-324-epic-mlx-first-class Sep 24, 2026
9 checks passed
@neuromechanist
neuromechanist deleted the feature/issue-351-phase15-final-parity branch September 24, 2026 06:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant