Skip to content

Epic #324: first-class MLX and reference-faithful fitting - #364

Merged
neuromechanist merged 27 commits into
devfrom
feature/issue-324-epic-mlx-first-class
Sep 24, 2026
Merged

neuromechanist merged 27 commits into
devfrom
feature/issue-324-epic-mlx-first-class

Conversation

@neuromechanist

@neuromechanist neuromechanist commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Epic #324 makes the MLX backend first-class end to end, and brings every backend's fitting closer to the Fortran reference.

Reviews and oracle tests along the way found several places where pamica's fit departed from amica15.f90. The epic grew to fix them, so default fits differ from 0.3.3 on every backend; the changelog's Unreleased section opens with the list.

Phases

Each phase was squash-merged into the epic after its review (pr-review-toolkit, Sonnet) and green CI.

Phase Issue PR Change
1 #323 #326 Shared pcakeep/pcadb policy; MLX support
2 #322 #325 (dev) AMICAICA.apply restores the PCA residual
3 #304 #327 One params-file reader for every backend
4 #313 #330 backend="mlx" in AMICA and AMICAICA
5 #306 #329 Guards on raw backend accessors
6 #315 #332 validate_implementations.py --backend all; end-to-end workflow test
7 #333 #340 doscaling normalizes component rows
8 #334 #342 Components stored as rows; correct share_comps
9 #335 #338 1-based schedule gates
10 #336 #337 Sphere written column-major
11 #339, #345 #348 Reference iteration order; unconditional A-freeze
12 #341 #350 Normalized initial mixing matrix
13 #344 #349 Single-precision density constants
14 #346 #347 AMICA_NumPy rejects unknown options
15 #351 #356 Every parity figure re-measured, including the paper's Table 1
16 #352 #353 Documentation audited against the code
17 #354 #355 One default set across entry points; raw PyTorch CPU fallback on Apple Silicon

Parity after the epic

Re-measured against the pinned v0.3.3 binary (.context/issue-351/findings.md).

Measurement Before the epic After the epic
ds002718, 5 seeds x 2000 iterations: LL gap to the reference within ~5e-4 6e-6
Same run: mean matched correlation (reference against itself 0.998) 0.998 0.9996
Harness, 100 iterations, from a shared start: LL gap 2.35e-4 1.6e-6
Newton, 2000 iterations on 70 channels, from a shared start: correlation 0.9974 0.9999998
Multi-model ensemble LL, pamica vs reference (Kolmogorov-Smirnov p) -3.3629 vs -3.3539 (6e-5) -3.3541 vs -3.3543 (0.83)

Per-iteration cost is unchanged.

Also in this PR

  • The paper's AI usage disclosure credits the Research Skills plugin collection (https://github.com/neuromechanist/research-skills) as the development harness.
  • CI:
    • The Python jobs skip a change that touches no code: docs, the paper and its rebuilt PDF, citation metadata, and non-Python .context/ records.
    • The rebuilt paper.pdf is committed back on dev and main only.

Follow-ups

Filed as their own issues: #328, #331, #358, #359, #360, #361, #362 and #363.

Closes #323, closes #304, closes #313, closes #306, closes #315, closes #333, closes #334, closes #335, closes #336, closes #339, closes #341, closes #343, closes #344, closes #345, closes #346, closes #351, closes #352, closes #354, closes #324.

neuromechanist and others added 27 commits September 22, 2026 14:09
* Add shared pcakeep/pcadb validation policy

validate_pca_reduction rejects non-integer or < 1 pcakeep and
non-finite or <= 0 pcadb; pca_reduction_requested answers the upfront
mir_step question; numerical_rank re-validates at fit time. Rank
policy unit tests pass (44).

* Validate pcakeep/pcadb in torch and NumPy

Both constructors call the shared validate_pca_reduction. The torch
mir_step gate delegates to pca_reduction_requested with the fitted
data's channel count, so pcakeep >= n_channels (the bundled
input.param's pcakeep 32) is no longer rejected. New regression test
fails before the fix and passes after; wrapper suite green.

* Persist torch pcakeep/pcadb as plain scalars

The shared validator accepts numpy integers, which the weights_only
torch.load in AMICA.load refuses to unpickle; state_dict now casts to
int/float. None stays None, so default payloads are unchanged.

* Add pcakeep/pcadb to the MLX backend

AMICAMLXNG takes pcakeep/pcadb with AMICATorchNG's position,
semantics and shared validation, passes them to numerical_rank, gains
the same upfront mir_step gate and message, and persists both in
state_dict config (additive; older payloads load with None). Docstrings
and the mir() error no longer claim MLX lacks explicit reduction. MLX
mir/persistence/backend suites pass (68).

* Add cross-backend pcakeep/pcadb tests

Real sample: torch/NumPy/MLX keep the same rank and sphere for
pcakeep=20 and pcadb=18.5 (rank 10); pcakeep wins over pcadb; invalid
values raise at construction; torch and MLX give the same mir_step
gate verdict; torch vs MLX reduced fits match (min |corr| > 0.999);
MLX and torch persistence round-trip the request. 61 passed; 53 of
them fail on the pre-#323 sources.

* Reword constructor comments flagged by typos

Comment-only; full-repo typos check passes.

* Document shared pcakeep/pcadb policy

Differences guide: backend-table row, at-a-glance row 11 (pcadb is a
pamica extension; pcakeep wins), a pcakeep/pcadb section, and the
mir_step gate section rewritten for both backends. MLX API page and
changelog entries. mkdocs build --strict passes.

* Ignore pcakeep/pcadb without sphering

No backend reduces when do_sphere=False (the reference sets numeigs =
nx there, amica15.f90:527), yet the mir_step gate still fired. The
shared predicate takes do_sphere and validates first; a new
log_ignored_pca_request (called by all three constructors) warns once
that the request is ignored, and carries the precedence INFO note, so
validate_pca_reduction is now validation only. Rank policy and
cross-backend tests pass (126).

* Refit and restart from the input geometry

_preprocess shrinks n_channels to the kept rank, and torch/MLX then
validated the next fit's X against it: a refit, or the second of
n_restarts under any rank reduction (explicit or automatic), raised
'X has 32 channels, model expects 20'. Both backends now keep the
constructor's count in _n_input_channels, validate X against it, and
reset n_channels/n_comps from it at the start of every _fit_once.
_load_params derives it from the restored sphere's width, so a reloaded
reduced model refits. 16 new cross-backend tests; the 12 torch/MLX
cases fail before the fix, NumPy is the passing control.

* Test reassignment, numpy scalars, torch payload

A pcakeep/pcadb reassigned after construction raises the validator's
ValueError at fit time on torch, NumPy and MLX (and at the mir_step
gate on torch/MLX); the validator rejects numpy.bool_ and accepts numpy
floats for pcadb; a torch state dict without the two config keys loads
with None and keeps its reduced sphere. Rank policy and cross-backend
suites: 160 passed.

* Cite torch symbols, not lines, in MLX core

Every torch_impl/core.py line citation in mlx_impl/core.py (about 50,
many already stale and six shifted by this PR) now names the symbol
(AMICATorchNG.<method> or pamica.torch_impl.core.<name>), which does
not rot. Fortran line citations are unchanged. Comment-only.

* Update docs for review findings

Differences guide: backend-table intro describes the raw MLX backend
accurately; PCA, MIR and PMI defined on first use; row 11 wording; the
pcakeep/pcadb and mir_step sections cover do_sphere=False and fit-time
checks. MLX API page and changelog use semantic line breaks and add the
do_sphere, refit/restart and invalid-payload-load entries. mkdocs build
--strict and typos pass.
* Add shared params-file reader with JSON aliases

read_params_file (issue #304) is the single params-file entry point
every backend should use: content-sniffs JSON vs Fortran input.param
text and returns canonical pamica keys either way, aliasing a JSON
file's schema spellings (min_grad_norm/max_decs/numrej/num_mix_comps/
share_int) via one table. A file with both an alias and its canonical
key raises ValueError.

writestep/do_history/histstep move from FORTRAN_UNSUPPORTED_KEYS to
translated (identity) keys: the reader translates what any backend
supports, not just the wrapper; the legacy NumPy backend already
implements periodic checkpointing under these names.

* Route AMICA.from_params_file through read_params_file

Replaces the wrapper's own inline JSON/Fortran content sniffing with
the shared reader (issue #304). Behavior change: sample_params.json's
own max_decs/min_grad_norm/share_int settings are now applied (as
maxdecs/min_nd/share_iter) instead of silently landing in the "not
applied" warning.

* Route NumPy backend through shared params-file reader

AMICA_NumPy(params_file=...) and load_default_params now read through
read_params_file (issue #304) instead of a bare json.load, so the
legacy backend accepts the literal Fortran input.param text too, not
just JSON. Canonical keys map to this backend's own spellings
(min_nd -> min_grad_norm, maxdecs -> max_decs, share_iter ->
share_int); files/data_dim/field_dim come off the same parsed dict
instead of a second file read. A params-file setting this backend
does not consume is named in one logger.warning.

from_json_file is renamed to from_params_file (matches the PyTorch
wrapper's name, now accepts both formats); no alias left behind.

pdftype != 0 now raises NotImplementedError at construction: this
backend implements only the generalized-Gaussian source density and
previously silently ignored the setting. numpy_impl/params.json's
bundled default pdftype changes 1 -> 0 to match; the value was never
read by the fit path, so no previously-passing fit's numerics change.

* Dedup validate_implementations.py's Fortran alias table

_FORTRAN_ALIASES is now derived from fortran_params.
PAMICA_KEY_TO_FORTRAN_KEY (the shared reader's single source of truth
for pamica-key -> Fortran-keyword) plus the one JSON-only share_int
addition, instead of a hand-maintained duplicate. load_sample_data
reads sample_params.json through read_params_file, so min_nd/maxdecs/
share_iter land in run_pytorch_amica's ng_kwargs via the plain
_NG_PARAMS filter; the max_decs special case is no longer needed.

* Document the shared params-file reader (issue #304)

docs/guides/validation.md: read_params_file, the JSON alias table, the
writestep/do_history/histstep move, and the NumPy backend's own
params-file/pdftype behavior.

docs/guides/amica-differences.md: the backend-differences table's
Fortran input.param reader row (NumPy now yes, MLX awaits epic #324
Phase 4) and the NumPy pdftype NotImplementedError.

docs/api/numpy-backend.md: the renamed from_params_file classmethod
and the pdftype guard.

docs/changelog.md: Unreleased entry covering the shared reader, the
wrapper's now-applied JSON aliases, NumPy's input.param support, the
from_json_file rename, and the NumPy pdftype guard.

* Route NumPy CLI through the shared params-file reader

cli.py's load_params() did a bare json.load, so the CLI died with
JSONDecodeError on a Fortran input.param even though AMICA(params_
file=...) already accepted it. It now calls the same
_read_numpy_keyed_params helper the constructor uses (no duplicated
key-mapping table), so both formats work; files stays unresolved
relative to the working directory, matching Fortran-binary semantics.

New test_numpy_cli.py exercises main() end to end on real bundled data
for both formats; verified to fail with JSONDecodeError on the
pre-fix load_params().

* Compose _FORTRAN_ALIASES from the shared alias tables

Replaces the hand-added share_int entry with a composition of
JSON_ALIAS_TO_CANONICAL (JSON alias -> canonical) and
PAMICA_KEY_TO_FORTRAN_KEY (canonical -> Fortran keyword), so
write_fortran_param_file accepts either raw JSON-schema keys or
canonical keys with no hand-maintained entry left in this module.

Adds coverage: canonical maxdecs/min_nd are written as Fortran
max_decs/min_grad_norm, and an end-to-end test feeds load_sample_
data()'s canonical-keyed params straight into write_fortran_param_
file, asserting the written file carries Fortran's own spellings and
none of the canonical ones.

* Add precedence, warning and table-size regression tests

NumPy: explicit kwargs win over params-file defaults (mirrors the
PyTorch wrapper's precedence test). Wrapper: fitting from input.param
names writestep/do_history/histstep in the "not applied" warning,
since AMICATorchNG has no matching constructor keyword (issue #312).
Pins FORTRAN_TO_PAMICA_KEY/FORTRAN_UNSUPPORTED_KEYS table sizes so a
future move between them forces the docstring/doc counts to be
revisited.

* Fix stale docs from PR #327 review

fortran_params.py module docstring and its mirror in validation.md
still claimed share_int was "out of scope" and surfaced only via the
"not applied" warning -- stale since read_params_file/
JSON_ALIAS_TO_CANONICAL landed; both now correctly say share_iter
needs no Fortran-side rename and the JSON alias table handles
share_int on the way in.

Reworded four "silently landed in the not applied warning" claims
(amica.py, changelog.md, validation.md, a test docstring): the
warning already named max_decs/min_grad_norm/share_int before this
epic, it just did not apply them -- "silently" overstated that.

numpy_impl/core.py's __init__ docstring now says params_file accepts
both formats. AGENTS.md's architecture-map line for fortran_params.py
now describes it as the shared reader for the wrapper, AMICA_NumPy
and the NumPy CLI. amica-differences.md spells out "generalized
Gaussian (GG)" at first use. Applies semantic line breaks to the
Markdown prose this epic added (changelog.md, validation.md,
amica-differences.md, numpy-backend.md), without reflowing
pre-existing text.

* Fix unparsable spelling flagged by typos
* Guard AMICATorchNG accessors on degenerate fits (#306)

transform/get_*matrix/get_rho/variance_order/model_loglik/
model_probability/mir/pmi now refuse a degenerate fit or a
non-finite parameter, and validate X shape like fit() does.
from_state_dict wraps a malformed config's TypeError as ValueError.

* Guard AMICAMLXNG accessors on degenerate fits (#306)

Same guard as AMICATorchNG (backend_parity.md): transform/
get_*matrix/get_rho/variance_order/model_loglik/model_probability/
mir/pmi refuse a degenerate fit or a non-finite parameter, and
validate X shape like fit() does. from_state_dict (and load(), which
delegates to it) wraps a malformed config's TypeError as ValueError.

* Guard AMICA_NumPy accessors on degenerate fits (#306)

Same guard, ported to the legacy backend's own converged/
stop_reason bookkeeping: transform/get_weights/
get_sensor_mixing_matrix refuse a degenerate fit or a non-finite
parameter, and transform validates data shape like fit() does.

* Add cross-backend backend-guard tests (#306)

Parametrized over torch, MLX (importorskip) and NumPy: every
guarded accessor raises on a degenerate stop_reason and on a
non-finite parameter, and still works on a healthy fit; every
data-taking accessor rejects a 1D or wrong-channel-count array,
including after a pcakeep reduction; from_state_dict raises
ValueError (chaining TypeError) on a malformed config.

* Document raw-backend degenerate-fit guard (#306)

amica-differences.md row 5: the raw torch/MLX/NumPy backends now
refuse degenerate models themselves, not just the AMICA wrapper.
changelog.md: note the behavior change under Unreleased.

* Add shared model_probability_from_loglik helper (#306)

Factors the NaN-vs--inf diagnosis and column-wise softmax
normalization out of AMICATorchNG/AMICAMLXNG.model_probability so
neither backend duplicates it; lives under pamica/metrics/,
matching the existing mir/pairwise_mi split (backend computes the
array, a shared pure function does the rest).

* Guard get_pdftype/shared_components, dedupe sweep (torch)

Reviewer follow-up on #306 (PR #329):
- get_pdftype/shared_components now go through _check_usable too.
- New _nonfinite_params() batches the isfinite sweep into one
  stacked reduction (one device sync instead of one per parameter,
  skipping None entries), used by _check_usable, state_dict and
  write_amica_output instead of each repeating the sweep.
- model_probability no longer runs the guard twice (its own
  _check_usable plus model_loglik's): the core Lht computation
  moved to _model_loglik_unchecked, called by both public methods
  after their own single guard + shape check.
- model_probability's NaN-vs--inf diagnosis now delegates to the
  shared pamica.metrics.model_probability_from_loglik.
- mir()'s Raises docstring now covers the degenerate-fit
  RuntimeError and the bad-shape ValueError.

* Guard get_pdftype/shared_components, dedupe sweep (MLX)

Same follow-up as the torch commit, ported to MLX:
- get_pdftype/shared_components now go through _check_usable too.
- New _nonfinite_params() batches the isfinite sweep into one
  stacked reduction -- one MLX graph evaluation instead of one per
  parameter (measured ~2.2ms -> ~0.4ms for the sweep, transform()
  on 512 samples ~2.5ms -> ~0.6ms), skipping None entries, used by
  _check_usable, state_dict and write_amica_output.
- model_probability no longer runs the guard twice: the core Lht
  computation moved to _model_loglik_unchecked, called by both
  public methods after their own single guard + shape check.
- model_probability delegates its NaN-vs--inf diagnosis to the
  shared pamica.metrics.model_probability_from_loglik.
- mir()'s Raises docstring now covers the degenerate-fit
  RuntimeError and the bad-shape ValueError.

* Name what get_weights returns in its guard message (#306)

Reviewer follow-up: the action string said 'get the unmixing
matrix', not what the public method itself is called.

* Extend backend-guard tests: pdftype, shared_components, pca (#306)

get_pdftype/shared_components added to the torch/MLX accessor
table, so the existing degenerate/non-finite/healthy parametrized
tests now cover them. The pcakeep channel-count test now loops
over every data-taking accessor (transform/model_loglik/
model_probability/pmi), not just transform; mir is deliberately
excluded there since it always rejects a pcakeep-reduced model on
other grounds (rank-deficient sphere).

* Test model_probability_from_loglik directly (#306)

Real Lht from a real 2-model fit's model_loglik() on the bundled
sample: a healthy Lht's columns sum to 1 and match the backend's
own model_probability() (behavior-preserving); a column poisoned
entirely to -inf and a single entry poisoned to NaN raise distinct
messages.

* List get_pdftype/shared_components in the changelog entry (#306)

The Unreleased entry for the backend-guard behavior change now
names the two accessors added in the PR #329 review round.
* test: pin loading a version-1 AMICA save

* feat: add float64 preprocessing accessors

* feat: select the AMICA backend (torch or mlx)

* feat: save the AMICA backend (format v2)

* feat: select the AMICAICA backend

* test: cover MLX export and torch-MLX agreement

* docs: document backend selection in the wrappers

* fix: name a malformed MLX save payload

* docs: note the MLX wrapper in write_amica_output

* fix: word the unexpected-keyword error for both

* feat: show the backend in AMICA and AMICAICA repr

* test: run MLX load errors without MLX

* test: degenerate AMICAICA fit builds no basis

* test: pin the weights-only save refusal

* test: drive MLX from the JSON params file

* test: cover fit_transform on both backends

* test: fold the v1 backend check into the pin

* docs: add Raises to the AMICA accessors

* docs: add backend Raises to AMICAICA

* test: AMICAICA propagates fit keyword errors

* docs: name both backends in the API index

* docs: point getting-started at backend=mlx

* docs: keep MLX precision figures in one place

* docs: mark the example's external input files

* feat: guard the preprocessing accessors (#306)
* feat: validate every backend in the harness

validate_implementations.py gains --backend {torch,numpy,mlx}, a comma-separated list, or all. Each backend runs the bundled sample with the harness's canonical params (NumPy through its own key map, MLX through AMICAMLXNG) against one shared reference run; an explicit --backend adds a parity summary table with runtime. The default torch-only run prints the same report as before (diffed).

Tested: dispatch and argument-parsing tests (no binary), and the AMICA_RUN_FORTRAN-gated per-backend parity test against the v0.3.3 native binary (3 passed).

* refactor: run the harness MLX row via AMICA

With Phase 4 merged, the torch and MLX runs share one AMICA(backend=...) path, so the MLX row exercises the wrapper a user calls. Numbers unchanged (same summary table as the direct AMICAMLXNG run).

Tested: harness dispatch tests (29 passed); all-backend run against the v0.3.3 native binary.

* docs: use American English spellings

35 British spellings in comments, docstrings and Markdown (behaviour, colour, neighbour, cancelled, licence, optimiser, ...). Proper nouns (Centre de Recherche...) and 'analogue' left as is; no identifiers touched.

Tested: ruff, typos.

* docs: cite torch core by symbol, not line

Seven torch_impl/core.py:<line> citations outside the MLX module pointed at drifted lines; they now name the method (or are dropped where the sentence already names it). Also replaces a stale pamica.py:802-813 citation (a file renamed in #34) in AMICATorchNG._newton_direction. Fortran line citations unchanged.

Tested: ruff.

* docs: MLX getting-started path and harness rows

getting-started gains a complete Apple Silicon (MLX) section (install, fit/transform/save/load, params file, AMICAICA, precision). The validation guide documents running the harness per backend and records the per-backend parity rows with their expected bars. The differences page records do_sphere=False (#328) and NumPy refit warm-start (#312), both verified against the code. AGENTS.md, changelog and the native-backend and testing pages follow.

Tested: mkdocs build --strict (scratch env); anchors checked in the built site.

* test: keep suite output out of the repository

An audit of the non-slow suite found 63 tests (13 files) writing ./output via AMICA_NumPy's default outdir; the slow CLI, end-to-end and pdf-family tests wrote test_output, _ng_e2e_tmp and scratch_amica_parity into the repo. A session-scoped autouse fixture now runs each session (per xdist worker) from a temporary directory, the harness resolves its bundled inputs from its own location, and the slow tests write under tmp_path. Stale .gitignore entries removed.

Tested: full non-slow suite (1335 passed, nothing written to the repo); coverage reports still land in the repo root with and without xdist.

* test: end-to-end workflow on torch and MLX

New mne_tests/test_end_to_end_workflow.py mirrors our own pipeline on the bundled EEG: average reference, AMICAICA with pcakeep = n_channels - 1 on each backend, get_sources vs transform, a single-component exclusion (exact back-projection, residual kept, rank drops by one), EEGLAB export reloaded by loadmodout with the padded sphere, AMICA save/load, an input.param-driven fit, and torch vs MLX Hungarian-matched sources >= 0.999. Not marked slow (about 5 s); MLX tests skip without MLX or an Apple GPU.

Tested: 12 passed locally (torch and MLX); measured tolerances recorded in the module.

* fix: transpose NumPy sensor mixing matrix

AMICA_NumPy.get_sensor_mixing_matrix returned pinv(sphere) @ stored A without the transpose the torch and MLX backends apply (the stored W is the unmixing transposed, issue #24), so its columns were rows of the true mixing matrix: about 10% off the torch maps after five iterations and no inverse of get_weights() @ sphere. New cross-backend test pins torch/NumPy agreement (4.6e-12 full rank, 9.5e-11 at pcakeep=20), MLX agreement at float32 tolerance, and the identity on every backend.

Tested: new test fails without the fix (4 NumPy cases) and passes with it (10 passed); backend-guard and share-comps suites green.

* chore: remove stale 2025 debug output

pytorch_debug_test/out.txt and pytorch_output_test/out.txt were run logs committed in August 2025 by the long-removed basic PyTorch backend; nothing reads them. The only reference, a typos exclusion for the first directory, goes with them.

Tested: typos.

* docs: pin the harness rows to the v0.3.3 binary

States why the rows name an explicit release binary instead of the resolver default (its cached 'latest' is never refreshed) and how to fetch it.

Tested: mkdocs build --strict.

* Rebuild paper.pdf

* fix: NumPy writes files only with an outdir

AMICA_NumPy defaulted outdir to ./output, so every fit wrote out.txt at construction, writestep checkpoints, history and its final results into the caller's working directory. The default is now outdir=None, which writes nothing (as torch and MLX never do unless asked); an explicit outdir (keyword, params file, or the CLI's --outdir, whose own default stays output) writes exactly as before. Behavior change recorded in the changelog and the NumPy API page; the conftest chdir stays as defense in depth.

Tested: new test_numpy_outdir.py (default leaves the working directory empty with every write path on; explicit and params-file outdir write as before; fails on the old code) and a CLI test that a run without --outdir still writes ./output; NumPy suites (107 passed).

* fix: NumPy viz plots the true maps and sources

plot_components drew stored-A columns (rows of the true mixing matrix) as mixing vectors, and it and plot_pdf_fits built activations as stored W times raw data (no transpose, sphere or mean); plot_model_comparison skipped the sphere. One helper, _component_maps_and_sources, now computes exactly what get_sensor_mixing_matrix and transform return, and all three plots use it. load_results reads a rank-reduced fit's zero-padded sphere instead of raising.

Tested: new test_numpy_viz.py compares the helper and the plotted lines/histograms with the live accessors, full rank and pcakeep=20 (10 passed; 8 fail on the old code).

* test: pin the default harness run and summary

Pins the literal default run (no --backend: one validation_report.txt, no parity_summary.md, exit 0), renders format_parity_summary with populated rows from two real runs (torch as the stand-in reference, NumPy through compare_results) and checks every numeric column, and runs the no-MLX exit over both 'torch,mlx' and 'all'.

Tested: 32 passed, 3 skipped (the AMICA_RUN_FORTRAN-gated parity test).

* docs: fix two more British spellings

unlabelled and relabelled, missed by the first sweep's word-boundary match. Test names containing 'honour' (identifiers) are left as is.

Tested: typos, ruff.

* docs: trim the getting-started MLX route

The MLX section duplicated the backends guide's end-to-end and precision sections; it is now install, AMICA(backend="mlx") fit/transform, an AMICAICA one-liner and the float32 note, linking to the guide for the full workflow. The install heading no longer uses a bare GPU abbreviation.

Tested: mkdocs build --strict; linked anchors checked.

* docs: harness status and test working directory

progress_summary.md's harness line now matches AGENTS.md (every backend, #315), and the testing guide explains the conftest session chdir and how a test resolves repository paths.

Tested: mkdocs build --strict.

* docs: one description of reference resolution

default_reference_binary's docstring is now the single full description of the reference-binary resolution order; the --native-engine help summarizes it and points there. The native-backend page runs the harness with uv run.

Tested: --help renders; harness tests (32 passed).

* test: two-iteration fits for the map gate

The NumPy-vs-torch sensor-map gate (1e-9) failed on the macOS arm64 runner at 1.3e-9: the backends start within 1e-13 and EM amplifies the gap about sixfold per iteration, so five iterations was platform-sensitive. Two iterations keep the same gate with a 1e-12 gap while the pre-fix transposed maps still differ by 2.6-3%. Tested locally: module passes.

---------

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Write and read the EEGLAB sphere column-major (#336)

* Fix S-order tests to expect column-major (#336)

* Add cross-backend sphere export round-trip test (#336)

* Document the sphere-order fix (#336)

* Add native-binary sphere orientation test

* Compare MLX sphere bytes bit-exactly

* Fix sphere-order comments in write_amicaout

* Warn about old asymmetric-sphere output

* Correct scratch-history note on S order

* Polish sphere-order wording in docs/tests
* Add seeded native-binary oracle helper

Writes a pamica state in the reference's load_* formats, runs the
pinned v0.3.3 binary and reads its raw output arrays back, for
element-wise oracle tests (issue #333). Used by the tests that follow.

* Normalize component rows in doscaling

Each model's stored A block is the transpose of the reference's
mixing matrix, so a component is a block row. doscaling normalized
stored columns, which is not a change of scale; torch, NumPy and MLX
now rescale each component row with the matching mu/beta, exactly as
amica15.f90:1843-1851 does (issue #333).

Tested: invariance, cross-backend and doscaling=False byte-identity
tests; seeded native oracle matches A/mu/sbeta to round-off. Tests
that encoded the column rule are updated.

* Retune overshoot recipes for the new trajectories

With components rescaled, the aggressive-Newton recipes run monotone,
so keep_best restores never fired (failures or silent skips). Add
newtrate=3.0 and longer budgets where needed, move the pdftype=1
switch to kurt_start=6, pick n_mix=1 for the default min_dll stop,
and check the keep_best recompute only after a restore.

Tested: every retuned test passes and no longer skips.

* Retune two MLX cross-backend test configs

The variance-order guard rejected the 30-iteration fit (gap 5e-4), so
run 50 (gap 2.4e-3). One two-model sample sat on the rejsig=2.0
threshold between float32 and float64; use rejsig=2.5.

Tested: both tests pass for one and two models.

* Count scalestep from 1 and validate it

The reference parses scalestep but rescales every iteration. Keep it
as a pamica extension counted from 1 (iterations s, 2s, ...), so the
default 1 is the reference, and reject a non-integer or < 1 value at
construction on every backend instead of dividing by zero mid-fit.
Record it as row 13 of the differences page.

Tested: cadence and validation tests on torch, NumPy and MLX.

* Document component orientation in ADR 0006

ADR 0006 states the storage convention (components are rows of each
stored block), the doscaling fix and the Phase 8 layout change, and
amends ADR 0001; the issue #24 notes point to it. The changelog
records the behavior change with the oracle numbers.

* Make keep_best-under-reject tests non-vacuous

Both the MLX and torch tests ran monotone fits, so final_ll_ equal to
the last LL proved nothing. Use the overshoot recipe, require that
it restores without do_reject and that the do_reject trajectory also
ends below an earlier peak, then assert no restore.

Tested: both tests pass; the old configs were monotone.

* Check reject agreement up to float32 rounding

Exact agreement of the rejected set is config luck: float32 and
float64 per-sample LLs differ by up to ~0.1 by the second pass.
Record each pass, require every disagreement to lie within the band
measured on that pass, cap the band and the disagreement count, and
add controls where a threshold or trajectory divergence fails.

Tested: seeds 7 and 8, one and two models, at rejsig 2.0.

* Note ADR 0003 figures predate the row rule

* Rebuild paper.pdf

* Derive MLX pin tolerances from float32 study

* Check variance order against a float32 band

* Band-check the pdftype1 reject test

* Use semantic line breaks in ADR 0006

* Note Phase 8 aligns the initial A norm

* Word pre-fix norm ranges as history

* Note parity rows predate the row rule

* Document when scalestep is validated

* Say oracle maxima span backends and models

* Mirror the zero-norm guard in Newton test

* Test the rescale on Newton-reached states

* Rescale a permuted comp_list across backends

* Test zero- and NaN-norm rows stay untouched

* Fail, not skip, on recipe preconditions

* Never read stale or partial oracle output

* Check out full history in pytest CI jobs

* Fail byte-identity pins in CI, not skip

* Use the CI-proven overshoot recipe in torch

---------

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Count schedule gates from 1 like the reference

Route every iteration-schedule gate in the torch, NumPy and MLX backends through a shared pamica/schedule.py. The Newton switch (iter .ge. newt_start), the numdecs reset on the switch-on iteration, the iter > newt_start ceiling ratchets, the outlier-rejection schedule and NumPy's restart-on-NaN window compared the 0-based loop index with the reference's 1-based setting and fired one iteration late. Share, kurtosis, A-freeze and writestep/histstep gates were already 1-based and now use the same helpers. Seeded against the v0.3.3 binary (newt_start=3, doscaling off, 6 iterations) the LL deviation drops from 1.2e-3 to 5.4e-11 (torch) and 2.1e-12 (NumPy).

* Update schedule-dependent tests to 1-based gates

Tests whose configurations or replays encoded the old 0-based newt_start/rejstart: the aggressive-Newton keep_best recipe moves from newt_start=1 to 2 (the exact pre-change trajectory its documented measurements came from), the LLt reject tests fire on iteration 6 of a 6-iteration fit, and the MLX ratchet/reset replays spell out the reference's 1-based arithmetic. Tested: the 11 affected files pass (179 passed, 1 skipped).

* Test schedule gates against 1-based iterations

Cross-backend real-fit tests (torch, NumPy, MLX): the first Newton step is iteration newt_start, the rho-rate ceiling ratchet opens only after it, the switch-on counter reset follows the reference's replay, rejection first fires on iteration rejstart, and NumPy's restart-on-NaN window covers the first restartiter iterations. Tested: 15 passed; against the pre-fix backends 10 of them fail.

* Add gated native oracle for the Newton start

Seeds the pinned v0.3.3 binary with pamica's own init (indir/load_*), newt_start=3, doscaling off, single-threaded, 6 iterations. Torch and NumPy match its LL trajectory and A to round-off (torch LL 5.4e-11, A 4.2e-10; NumPy LL 2.1e-12, A 1.1e-10) where the pre-fix code deviated by 1.25e-3 and 3.1e-2. A no-Newton control shows the compared window is Newton-sensitive. Tested with AMICA_RUN_FORTRAN=1: 3 passed; 2 fail on the pre-fix backends.

* Document the 1-based schedule gates

Changelog entry for issue #335: which gates moved, the native-oracle numbers before and after, the default-path cases that change, and how to reproduce an old trajectory. No row in docs/guides/amica-differences.md changes: the anchored share A-freeze is already recorded there.

* Rebuild paper.pdf

* Validate iteration-schedule settings

One shared validator, schedule.validate_iteration_setting, called by all three constructors with identical messages: newt_start must be an integer >= 0 whether or not do_newton is on (it also gates the rho-rate ratchet), and rejstart an integer >= 1 when do_reject is on (with 1-based counting, rejstart <= 0 skipped the reference's unconditional first pass; NumPy accepted 0, torch and MLX any value). NumPy also validates restartiter >= 0, maxrestarts >= 0 and histstep >= 1 under do_history (histstep=0 was a bare ZeroDivisionError mid-fit), and documents that restartiter=0 disables restart-on-NaN. schedule.every and periodic_due raise ValueError for a non-positive interval. Docstrings: newton_active cites amica15.f90:1666, newt_start=0 and 1 are stated to fit identically, and the MLX newt_start/newtrate entries document the newtrate ratchet. Tested: the new validation, restart-window and interval tests pass (36 passed with the reject suites).

* Reflow MLX schedule docstrings

The share_start/share_iter and kurt_start entries no longer break mid-clause. No code change.

* Record NumPy restart-on-NaN difference

Row 14 of the differences guide: the reference allows maxrestarts+1 restarts and, because startover is never cleared (amica15.f90:1046), never calls update_params again after the first one (:1115-1122); the NumPy backend allows maxrestarts restarts and resumes fitting. PyTorch and MLX have no restart-on-NaN path and stop with nan_ll.

* Correct the schedule changelog entry

Use the measured pre-fix deviations (1.25e-3 over 6 iterations, 1.33e-3 over 100) in the changelog and the oracle test docstring, list restartiter among the settings that now count from 1, and add the new validation.

* Test schedule gates at their boundaries

The arithmetic test gains the =1 rows for every helper (newton_switches_on at 1 and 0, past_newton_start at newt_start and newt_start+1, rejection with rejstart=1/rejint=1 through maxrej exhaustion, every/periodic_due at 1, the restart window at 1 and 0). The first-Newton-step test runs at newt_start 1 and 5, and a new real-fit test on every backend shows newt_start=0 and 1 give bit-identical trajectories on an overshooting run whose ratchets fire. Tested: 12 passed.

* Make the switch-on reset test discriminating

The replay-based reset tests passed with the reset one iteration late too. The cross-backend test now uses a searched configuration (4096 frames, lrate=0.8, newt_ramp=1, maxdecs=3, newt_start=3) whose likelihood decreases on iterations 2, 3 and 4, where the reference rule predicts no ratchet and the one-late rule predicts one; the fitted lrate ceiling must match the reference. It fails on all three backends with only newton_switches_on reverted and on the pre-fix backends. The MLX twin, which could not discriminate, is removed in favor of it. Tested: 3 passed; 3 failed with the reset reverted.

* Test the rho-rate gate at newt_start=20

A searched natural-gradient configuration (4096 frames, seed 4, lrate 0.5, maxdecs 2) completes a maxdecs cycle on iteration 21 on all three backends, so the rho-rate ratchet gate is now pinned at the shipped default: open for newt_start=20, closed for 21. This is the one default-path change of issue #335. Tested: 3 passed; 3 fail on the pre-fix backends.

* Document the reset-test configuration search

States the full search that selected the switch-on reset configuration, including the two-model hit found after the first commit.

* Shift Phase 7 recipes to 1-based settings

Phase 7 retuned its overshoot recipes under 0-based gating. With newt_start and rejstart counted from 1, n becomes n+1 so each documented trajectory reproduces bit for bit: the min_dll overshoot recipe (newt_start 2, newtrate 3.0) again peaks at 70 and stops at 71 on torch (1.9e-4 below the peak) and at 63/64 on MLX (4.2e-4); its do_reject variants move rejstart 5 to 6; the doscaling-rows NEWTON state uses newt_start=2; the min_dll-reachable test uses 6 and the slow rholrate ratchet test 51. Tested: the affected suites pass.

* Route scalestep through the schedule module

Phase 7's inline (iteration + 1) % scalestep rule becomes schedule.every in all three backends, and its scalestep check becomes schedule.validate_iteration_setting (same semantics and message; still only when doscaling is on), which drops the numbers imports it needed. Tested: test_doscaling_rows passes.

* Re-search the newt_start=20 gate configuration

Phase 7's component-row doscaling moved the old configuration's second ratchet off iteration 21, and the test's precondition guard failed as intended. The re-searched configuration (4096 frames, seed 1, lrate 0.6, maxdecs 3, newt_ramp 1) decreases at ll indices 4, 5, 8 and 9, 19, 20 on all three backends (smallest decrease 4.7e-4), ratcheting at 8 and 20. Tested: 3 passed.

* Note the combined doscaling and Newton-start parity

Changelog: with doscaling on (component rows, issue #333) the seeded 100-iteration Newton comparison stays within 3.9e-6 (torch) and 7.0e-6 (NumPy), where the #333 fix alone left 2.1e-4; scalestep now uses the shared schedule helper.

---------

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Add a loader for pre-change pamica in tests

pamica/tests/pre_change.py imports the whole package at a pinned commit
from git under a private name, so byte-identity tests can compare the live
backends with the pre-change ones in one process. It fetches only when the
object is missing (a --depth 1 fetch in a full clone makes it shallow), and
fails under CI but skips locally when the commit is unreachable.

* Store mixing matrix components as rows

All three backends now hold A as (n_comps, n_channels): row k is the
reference's column A(:, k), and model h's block is A[comp_list[:, h], :],
the same matrix as before, so W = inv(block) is unchanged (issue #334).
comp_list now indexes components in A as it does in the densities, so:

- the share metric compares pinv(sphere) @ A.T, the components' scalp
  maps, not stored columns of each block;
- a merge ties two components' mixing vectors and densities, and the
  gm-weighted dAk/zeta step averages one component across models;
- doscaling rescales each component row once, shared or not;
- accessors, W, the EEGLAB export (reference layout), load_results and
  the viz helpers follow the new layout.

Fits without a merge are byte-identical on every backend; ndtmpsum moves
by float round-off (it sums per component row, like the reference).
Existing sharing, doscaling and MLX tests are updated to the row layout;
two short sharing recipes are retuned because early scans now merge
most components. Tested: full non-slow suite and slow tests.

* Version saved models for the row layout

The PyTorch state_dict is now format_version 4 and the MLX save format 2.
A version 3 (PyTorch) or 1 (MLX) payload is converted through the shared
pamica.component_layout.rows_from_legacy_columns: without loss when its
comp_list is unmerged, and refused with a request to refit when a
component id is shared across models, since that merge was made under
the old column semantics. AMICA.save files (still version 2) go through
the same path.

Tested with test_component_rows_persistence.py (24 passed): round trips
in the new formats, and payloads written by the pre-change classes
loaded from git, converted or refused.

* Refuse a stale multi-model A in load_results

load_results now checks that each model's block of A inverts the W
written beside it (max |A_h W_h - I| <= 1e-3; measured at most 3.8e-7
for float32 MLX exports and 1.1e-15 for float64). A multi-model
directory written before issue #334 stored A with components as columns
and reads back scrambled (residual 1.5); it is refused with the remedy
instead of handing the viz helpers wrong sensor maps. Single-model
files are the same bytes in both layouts and still load.

* Test component rows and sharing semantics

test_component_rows.py pins, on the sample EEG:
- byte identity of every unshared configuration against the pre-change
  package (3627a6e) on PyTorch, NumPy and MLX: 1-3 models, doscaling,
  Newton, pdftype 0/1 and do_reject;
- the metric compares get_sensor_mixing_matrix columns, a planted
  duplicate merges, and a merged group shares one mixing vector;
- the EEGLAB A file is the reference layout, single-model exports are
  byte-identical, and load_results refuses a pre-change multi-model A;
- opt-in (AMICA_RUN_FORTRAN=1): updates from a merged load_comp_list
  state match amica15mac to float64 round-off on PyTorch and NumPy,
  with doscaling on and off.
All pass locally, including the gated oracle.

* Document the component-row layout

ADR 0007 records the row layout, the sharing semantics that follow from
it, the persistence and export policy, and the measured evidence; ADR
0006 and 0001 are amended, and ADR 0006's initial-A normalization line
now points to issue #341. The changelog, the differences page (sharing
rows, a new section for issue #334, and row 15: the reference's
single-precision density normalizers), and the backends, EEGLAB,
validation and API pages describe the new behavior. mkdocs build
--strict passes.

* Reflow two sharing test docstrings

Docstring text only; no test logic changes.

* Skip MLX wrapper routes without MLX

The merged-save refusal test's wrapper route fitted an MLX model without first checking MLX was installed, so it failed on the Linux runner. _wrapper_fit now performs the check for every wrapper route. Tested with MLX hidden: the new test files pass or skip.

* Pin byte identity to the Phase 9 epic head

The pre-change package is now 0930c0e, the epic head with Phase 9's schedule gates and without this phase, so the two sides differ by the layout change alone. The full byte-identity matrix (Newton included) and the persistence tests pass against it: 92 passed, 2 gated oracle tests skipped.

* Stop tests from fetching pinned commits

Tests that load historical code ran `git fetch origin <sha> --depth 1`,
which in a full clone records a shallow boundary that gc can prune
behind (issue #343). pre_change.py now only reads: it checks the pin
with `git cat-file -e` and, when the commit is missing, fails under CI
and otherwise skips with the command that fetches it (`git fetch origin
<sha>`, or `git fetch --unshallow origin` in a shallow clone). Pins must
be full SHAs.

test_mlx_fit_noop.py, the last loader that fetched, now imports the
whole package at its pin through the same helper; its fits are still
bit-identical. test_pre_change_loader.py runs the loader behind a git
wrapper that logs its arguments and execs the real git, and asserts no
fetch ran and the shallow state is unchanged (a negative control with a
real fetch trips the assertion). Tested: loader and fit-noop tests pass,
also with CI set.

* Record the single-precision density constants

The reference writes its density normalizers as dble(<literal>), which
gfortran reads in single precision: log(dble(1.772453851)) in the
rho == 2 branch is 3.0e-8 above log(sqrt(pi)) (measured in the seeded
oracle), and the Gaussian and cosh families' normalizers differ from
pamica's double values by 3.7e-10, 2.0e-8 and -2.1e-8 (computed).
Row 16 of the differences page records this, citing issue #344, which
decides whether to adopt the reference's values. The torch and MLX
comments that called _LOG_SQRT_2PI and _LOG_NORM_COSH_* bit-for-bit
with the binary are corrected, and the merged-state oracle's
maxrho=1.99 cites the issue. Comments and docs only.

* Show early mass merges happen in the reference

With share_comps on, the pinned binary's scan runs but never merges:
Spinv2 is never allocated, so every similarity is NaN, even at
comp_thresh=0. Its similarity (Spinv2 = Spinv^T Spinv) applied to its
own state after 8 iterations from pamica's initialization merges exactly
the pairs the torch and NumPy scans merge (32, 32 and 30 at 0.9, 0.95
and 0.99). In the recipe that went non-finite, the reference's update
from pamica's post-scan states collapses in step: the second model's
gm falls to 5.04e-3 on both after the first scan, and after the second
reaches zero in one iteration and NaN in the next.

Two gated tests pin this (AMICA_RUN_FORTRAN=1, both pass). The
differences page, ADR 0007, the changelog, the validation table and the
torch/MLX comments that called the scan unrunnable now say what it does.

* Harden the pre-change loader and its guard

pre_change.py now fails under CI, and otherwise skips, with its own
message when git is not installed (instead of a raw FileNotFoundError)
and when the tests do not run from a git checkout (instead of advising
a fetch that cannot help). The repository root is a parameter, and a
cached alias must map to the same commit or loading raises.

The test's PATH wrapper now refuses fetch, pull and every depth or
shallow option without running git, so the committed negative control
(a fetch sent through _git) is caught by _assert_no_fetch and leaves the
clone untouched. New tests cover missing git, a non-checkout root and
alias reuse. Tested: 10 passed, also with CI set; the clone is not
shallow afterward.

* Validate comp_list dtype in the layout conversion

rows_from_legacy_columns now refuses a non-integer comp_list (float or
NaN ids) with its own ValueError before the bounds check, instead of
letting it reach the indexing. test_component_layout.py adds direct
unit tests on small arrays for every ValueError branch: the comp_list
rank, the dtype, the A shape, an out-of-range id, an id repeated within
a model, and an id shared across models (the refit message), plus a
block-for-block conversion of an unmerged, permuted comp_list.
Tested: 11 new tests pass; persistence tests unchanged.

* Refuse a truncated A in load_results

load_results now checks that the A file holds num_comps * nw values
before reshaping it to (num_comps, nw), and names the file and both
counts when it does not. Before, a length no reshape accepts surfaced
as a bare NumPy reshape error, and one short by a whole component
reshaped to the wrong width and failed later as a matrix-product error.
Tested on truncated copies of a real two-model export (both cases).

* Raise instead of assert in NumPy _write_results

The fitted-A and outdir preconditions of AMICA_NumPy._write_results are
now explicit RuntimeErrors, since assert is stripped under python -O.
The comment above them also said the reference omits A from its
output; the binary does write it, in the layout pamica now writes.
Tested: NumPy write and load tests pass.

* Widen the byte-identity matrix

The unshared byte-identity test now also covers the Gaussian, logistic
and sub-Gaussian cosh families (pdftype 2, 3 and 4, single model, n_mix
1 for family 4), four blocks of 1024 samples, a keep_best fit whose
log-likelihood falls so the safeguard restores an earlier snapshot
(asserted), and best-of-two restarts, on every backend that supports
each. The NumPy backend has only the generalized Gaussian and no
keep_best, so those cases run on PyTorch and MLX; the test says so and
test_numpy_rejects_other_pdftypes pins the first reason. The helper now
lets a configuration override its defaults, and compares final_ll_.
Tested with CI=1: 66 passed.

* Check the MLX double conversion on save again

test_old_mlx_save_converts_losslessly now saves the converted model
again, checks the file is format 2 with a component-row A, reloads it,
and compares A and every output with the saved model, as the torch test
does: a second load must not convert the already-converted A again.
Tested: both recipes pass.

* Fix the Spinv citation and sharing comments

The torch _identify_shared_comps docstring cited Spinv's allocation at :551; it is amica15.f90:569. A test_mlx_newton docstring said merged-away columns and curvature not indexed by mixing column; with components as rows the words are component and mixing vector. Comments only.

* Label the merged-oracle figures consistently

ADR 0007 and the changelog now say the 3-iteration oracle figures are the worst of doscaling on and off, and give the old column semantics' error as the matching worst value (A 0.21, LL 4.3e-4, not 0.20 and 4e-4); the test comment lists both measurements. Docs and comments only.

* Describe component-row sharing in the notes

AGENTS.md and .context/progress_summary.md described the old column metric and merge as working. They now describe the component-row semantics (ADR 0007, #334): the metric compares de-sphered component mixing vectors, the merge ties components, and the reference binary's own scan never merges because Spinv2 is never allocated. Docs only.
* Reject unknown AMICA_NumPy keyword arguments

AMICA_NumPy(**kwargs) forwarded every keyword into a params dict read
with params.get(...), so a typo (max_iters=50) or an option the legacy
backend does not implement (keep_best=True) constructed silently and
had no effect.

__init__ now validates **kwargs before any other processing: an
unrecognized name raises TypeError with a difflib-based suggestion,
and a name implemented on the PyTorch backend (AMICATorchNG) but not
this one raises TypeError naming it and pointing to
AMICA(backend='torch'). The accepted-keyword set is derived from
_CONSUMED_KEYS/_CANONICAL_TO_NUMPY_KEY (the same source the
params-file routing uses) plus this constructor's own named
parameters, so the two surfaces cannot drift apart. The canonical
spelling of the three renamed settings (min_nd/maxdecs/share_iter) is
now also accepted as a keyword argument directly, translated the same
way a params file's setting already is; passing both spellings at
once raises TypeError.

* Add tests for AMICA_NumPy kwargs validation

Real construction only, no mocks (no .fit() needed to exercise
__init__'s validation): a typo raises TypeError with a suggestion; an
unsupported cross-backend option raises the specific message; the
three canonical-spelling aliases are accepted and their conflict is
rejected; every documented NumPy option constructs, alone and all
together; and a source scan proves _CONSUMED_KEYS matches every
params.get(...)/raw_params.get(...) read site exactly, so the
accepted-kwargs set cannot silently drift from what __init__ actually
reads.

* Document AMICA_NumPy kwargs rejection behavior change

docs/changelog.md: Unreleased entry noting the legacy NumPy backend
behavior change (unknown or unsupported options now raise instead of
being ignored). docs/api/numpy-backend.md: a short paragraph mirroring
the existing pdftype callout.

* Improve AMICA_NumPy kwarg validation (review)

PR #347 review follow-up, several items together since they touch the
same functions:

- files/data_dim/field_dim now work as constructor keywords, with the
  same meaning as in params_file, instead of passing the accepted-
  kwargs check but staying silently inert (fit() with no data still
  said "no 'files' configured" even after a caller set one via a
  kwarg). Conflicting values between params_file and a kwarg raise;
  the same value from both does not.
- n_models/n_mix (AMICATorchNG's spelling of this backend's own
  num_models/num_mix) get a message naming the correct spelling
  instead of the generic (and here wrong) torch-only message.
- n_channels gets its own message: it is inferred from the data
  passed to fit() on every backend, not a constructor keyword on any
  of them (the AMICA wrapper's own fit(**kwargs) excludes it too).
- Every offending keyword in one call is named in a single error,
  however many different categories are mixed together. A first
  version raised on the first category found, so
  AMICA_NumPy(n_models=2, totally_bogus=1) named only n_models and
  dropped totally_bogus.
- _torch_only_options() imports AMICATorchNG lazily, inside its own
  body, called only on the error path, so numpy_impl carries no
  module-level dependency on torch_impl.
- Groups difflib/inspect with the other stdlib imports.
- Hardens the source-scan test: fails loudly on a params[...] bracket
  read or a non-literal params.get(...) key, either of which would be
  invisible to the scan.
- Adds typo-suggestion tests for params_file/use_tqdm alongside verbose.

* Fix dropped keyword in AMICA.fit mlx check

PR #347 review item 4: the same single-category drop
AMICA_NumPy._reject_unknown_kwargs had (issue #346) existed here too --
backend='mlx' with fit(dtype=..., blocksize=...) named only 'dtype'
(ValueError) and silently dropped 'blocksize' from the message, since
the torch-only check raised before the unknown-keyword check ever ran.

Both categories are computed before either is raised; a combination
now names both in one TypeError. The two single-category messages are
unchanged (still ValueError for torch-only alone, TypeError for
unknown alone), so existing callers/tests keep working.

* Document NumPy kwargs review follow-up

docs/changelog.md: a Review follow-up sub-bullet under the Phase 14
entry. docs/api/numpy-backend.md: mentions files/data_dim/field_dim as
keyword arguments and the n_models/n_mix/n_channels-specific
messages.
* Split the M-step into direction and update

Compute the natural-gradient/Newton step and its norm (the reference's
dAk and ndtmpsum) in _update_direction and apply it in
_update_parameters, in all three backends, sharing one set of
intermediates. The fit loops call both back to back, so every fit is
byte-identical (checked on 8 configurations per backend).

* Follow the reference's iteration order

Run the likelihood-decrease response and the stopping checks before the
parameter update, and exit on a stop before any parameter moves, as
amica15.f90:1015-1122 does, in all three backends. The step taken right
after a decrease now uses the halved rate, and a stopped fit returns the
parameters whose likelihood is ll_history[-1]. Keep the reference's
working rho rate and its ceiling apart (rholrate, rholrate_cap), reset
inside the A-update branch. Fits without decreases or stops are
byte-identical; seeded oracle deviations drop from 1.4e-2 to 2.7e-5.

* Hold A on the reference's schedule without sharing

Apply the reference's A-update guard (iter >= share_start and
mod(iter, share_iter) <= 5, amica15.f90:1803) in every fit, share_comps
on or off, through schedule.share_freeze, replacing the anchored
post-merge window. Every backend now rejects share_iter < 7 (NumPy:
share_int) whether or not sharing is on, since a shorter cycle would
never update A again. Default fits of 100 or more iterations change.

* Add iteration-order tests and native oracle

Always-on cross-backend tests for the freeze iterations, the decrease
timing, stop semantics, keep_best pairing and share_iter validation, plus
an opt-in comparison against the native binary (AMICA_RUN_FORTRAN=1).
All pass; the native comparison fails on the pre-change code.

* Retune tests to the reference's iteration order

Existing tests that pinned the old order (update before the checks, a
freeze gated on share_comps and anchored on share_start, one rho rate) now
pin the reference's, or use recipes that still exercise what they test.
Full non-slow suite and the gated native oracles pass.

* Document the reference's iteration order

ADR 0008, changelog entry with oracle and parity numbers, and the
differences guide: the A-freeze is now the reference's arithmetic, and
convergence stops return the parameters final_ll_ describes.
mkdocs build --strict passes.

* Test the rho-rate ceiling as its own value

Rate tests that read rholrate as the ceiling now read rholrate_cap, the
cross-backend state copies carry it, and a new test round-trips it through
state_dict on PyTorch and MLX, including a save without the key.
Affected tests pass, including the slow torch ratchet test.

* Stop overshoot recipes on the first decrease

The keep_best overshoot recipes stopped via min_dll=1e-4/maxincs=2,
which under the reference order ended at the peak or below it depending on
round-off: at the peak on the macOS CI runner, and in 1 to 2 of 12 data
perturbations locally. maxincs=0 with min_dll=1e-8 stops them on their
first likelihood decrease instead, which held in all 12 perturbations on
PyTorch and MLX, with and without rejection. Affected tests pass.

* Stop identically on non-finite values

Every backend now keeps a non-finite likelihood out of its history (NumPy
recorded it), stops before applying a non-finite step or gradient norm
(new degenerate stop_reason nan_direction), and stops on non-finite
parameters right after an update (nan_params, now also in PyTorch and
NumPy). NumPy defers only an A/W corruption inside its restart window to
restart-on-NaN, which can repair it. Cross-backend injection tests pass.

* Clear numincs on a NumPy restart

The reference's min_dll comparison with the NaN likelihood of its restart
iteration is false, so it zeroes numincs (numdecs is left as it was).
NumPy's restart carried the small-gain count across; it now clears it.
The new test stops at index 8 as expected and at 6 without the fix.

* Validate saved learning rates on load

from_state_dict (and the MLX file load) now refuses a missing or
non-finite lrate, lrate_cap, newtrate, rholrate or rholrate_cap with a
ValueError naming the field; a missing rholrate no longer surfaces as a
bare KeyError on PyTorch. Tested on PyTorch and MLX.

* Validate share_start on every fit

The A-freeze reads share_start whether or not share_comps is on, so every
backend now requires an integer share_start >= 1 always. The share_iter
message now says why a value below 7 is refused (A held permanently from
share_start), and NumPy's names both of its spellings. Both checks are
reached through a Fortran params file on all backends; tests pass.

* Test every convergence stop's parameters

The stop test now runs min_dll, grad_norm, grad_norm_floor and lrate_floor
on every backend, each with a precondition on its stop_reason and a twin
that runs the same iterations to max_iter. All 12 cases pass.

* Test that a NumPy restart skips the update

A pass-through recorder on _update_parameters shows the restarting
iteration applies no update; the same fit of the pre-change code, loaded
through pre_change.py as a control, updated on it. Passes.

* Test MNE n_iter_ after each kind of stop

to_mne_ica().n_iter_ equals iteration + 1 and len(ll_history) after a
min_dll stop and after a max_iter fit, on PyTorch and MLX. Passes.

* Take the oracle floor over two thread pairs

The reference noise floor is now the larger deviation of a 2- or 4-thread
binary run from the 1-thread run. The multi-threaded runs are not
reproducible, so the docstring gives each floor's range over three runs.
All four gated oracle cases pass in each run.

* Fix docstrings and Fortran citations

mir_history_ records no waypoint on a stopping iteration; MLX documents
its rholrate/rholrate_cap pair as PyTorch does; NumPy fit() gets the
iteration-order notes. Rejection is cited as amica15.f90:1136-1140 and the
check ranges start at 1056 and 1019. Docstring and comment changes only.

* Document the non-finite stops and validation

ADR 0008 states the final_ll_ invariant with its keep_best qualifier and
the review's changes; the changelog and differences guide cover the
non-finite stops, share_start validation and saved-rate checks, with the
two-pair oracle floors. AGENTS.md and the context notes describe the
unconditional A-freeze. mkdocs build --strict passes.
* Use the reference's single-precision constants

The reference writes its density normalizers as default-kind literals
widened with dble, so the binary uses their float32 roundings. All three
backends now read those values, and epsdble's, from one shared module,
pamica/reference_constants.py; MLX holds the rho == 2 normalizer in its
lgamma table. The literal-formula tests and the benchmark's copy now use
the reference's values. Tested: pdf-family and backend test files pass.

* Compare pre-change fits at the old constants

The byte-identity tests against code from before issue #344 (component
rows, doscaling off, the MLX pre-Phase-3 tip) now give the live backend
its old density constants through pre_change.use_pre_344_constants, so
they still isolate the change each was written for: their two-model and
non-GG fits reach the changed normalizers. Tested: the three files pass.

* Pin the constants against the reference

test_reference_constants.py checks each shared constant against the
float32 rounding of its literal on the cited reference line (exact
rational arithmetic), pins the sweep of amica15.f90 and its header,
checks that no backend keeps a copy, checks every backend's rho == 2
log-density, and (opt-in) seeds the native binary for pdftype 2, 4 and
1. Tested: 23 passed with AMICA_RUN_FORTRAN=1.

* Run the merged oracle at rho == 2

The merged-state oracle no longer holds maxrho at 1.99: its warm seed
now has mixtures clamped at 2, it asserts that, and the code before
issue #344 is replayed from the same state as a control (off by 2.8e-9
in log-likelihood after 1 iteration, against 2.2e-15 now). Tolerances
unchanged. Tested with AMICA_RUN_FORTRAN=1.

* Document the single-precision constants

Changelog entry with the oracle numbers; the differences guide gains a
section on the constants, and row 16 now records the one kind of
single-precision literal pamica keeps at its decimal value (defaults of
input.param keys); the validation guide notes the new family oracle.
Tested: mkdocs build --strict, typos.

* Note the reference's unset lrate0 in row 16

Without an lrate key the reference never sets its ceiling lrate0, so
its lrate drops to 0 after the first iteration (checked against the
pinned binary). Row 16 now says so and uses comp_thresh as the example.
Tested: mkdocs build --strict, typos.

* Skip MLX before patching its constants

On a host without MLX the byte-identity tests imported the MLX module
to patch it before the check that skips them. Resolve the classes
first. Tested: the two files pass with MLX and skip its cases without.

* Remove the unused compute_log_pdf

numpy_impl.pdf.compute_log_pdf had no callers in the package, tests or
docs. Tested: the pdf and viz tests pass.

* Plot the fit's density in compute_pdf

compute_pdf, which viz.plot_pdf_fits draws, now takes its normalizers
from pamica.reference_constants, including the reference's
single-precision sqrt(pi) at rho == 2, so it draws the density the fit
uses. Its docstring says it is a plotting helper with the legacy
pdftype numbering. Tested: the pdf, viz and constants tests pass.

* Scan every backend module for copies

The no-copy test now scans every .py file of torch_impl, numpy_impl and
mlx_impl with comments and strings blanked, and flags the linear forms
(sqrt(pi), sqrt(2*pi), ...) as well as the log forms and the literals.
A new test checks that it flags every definition this change removed.
Tested: passes; fails on the old pdf.py sqrt(pi) line (checked by hand).

* Assert the family held in the oracle

The family oracle now asserts that every source kept its family and no
kurtosis switch ran (n_kurt_done == 0), for pdftype=1 with num_kurt=0
too, in the live and the pre-change runs. Tested with
AMICA_RUN_FORTRAN=1.

* Check the pre-#344 constant substitution

A control test: every constant use_pre_344_constants patches differs
from the live value, the patch takes effect, it covers every changed
constant each backend module reads, and on MLX it restores the
pre-change rho == 2 table entry. Tested: passes on all three backends.

* Note the plotting helper in the changelog

compute_pdf now draws the fit's density and compute_log_pdf is gone.
Tested: mkdocs build --strict, typos.
* Run the native binary from its own draw

run_drawn_reference runs the binary with nothing loaded and a fixed
seed; with max_iter=0 it writes its drawn initialization. The run and
read-back are shared with run_seeded_reference.

* Normalize the drawn initial mixing matrix

Every backend now draws its initial A through the shared
pamica.initialization.initial_mixing: per model, 0.01*(0.5-u) with the
diagonal set to one and each component divided by its norm, as the
reference does (amica15.f90:805-823), including the NumPy restart
redraw. A supplied or loaded A is used as is. Tested in later commits.

* Start pre-change classes from the new init

The byte-identity pins (fb13d76, 0930c0e, 2e04006) and the #339 oracle
control now fit the historical class from the live normalized initial A
through pre_change.with_normalized_initial_mixing, so each still isolates
its own change. Tested: all those byte-identity tests pass.

* Test the normalized initial mixing matrix

Always-on cross-backend start checks with a pre-change control, the
supplied/loaded A paths, the NumPy restart redraw, and two gated native
oracles. Tested: 34 passed with AMICA_RUN_FORTRAN=1.

* Re-search the MLX pdftype=1 restore recipe

The normalized initial A made seed 3 at lrate 0.5 climb monotonically
to max_iter, so the restore was never exercised. Seed 6 at lrate 0.4
(kurt_start 7) overshoots by 0.17 after all five switch passes, on all
12 perturbed inputs. Tested: test_mlx_keepbest.py passes.

* Re-record the MLX fit-path canary

The normalized initial A moved the recorded trajectory (ll_history[0]
by 1.55e-5). Pins re-recorded on the M4 Pro and tolerances re-derived
from the same 48-draw float32 perturbation study, floored at one float32
unit: no draw moved ll_history[0]. Tested: the canary passes.

* Seed the doscaling oracle from a conditioned state

From seed 42's normalized init, one sample sits 2.3e-7 from a mixture
center; the mu update amplifies the 1e-14 sphere round-off to 1.3e-9
(PyTorch and NumPy differ by 7.9e-10), over the 1e-9 bound. Seed 45
stays 7x under every bound; a new precondition asserts the backends'
own disagreement sits 3x below each bound. Tested with the binary.

* Re-search the early-merge collapse recipe

From its normalized initial A, seed 20's second model keeps gm 1.2e-3
after the second scan instead of collapsing to zero in one iteration.
Seed 23 (of 23, 32, 40, 41 in seeds 0-49) collapses as the test needs.
Tested with the binary (AMICA_RUN_FORTRAN=1).

* Document the normalized initial mixing matrix

Changelog entry with the measured behavior change, ADR 0006's remaining
difference marked resolved (and ADR 0007's pointer), the differences
guide's init wording and the re-searched collapse recipe. mkdocs build
--strict passes.

* Re-seed the MLX MIR restore test

From its normalized initial A, seed 0's discarded step moved the MIR by
2.1e-4 relative (3.4e-5 to 2.1e-3 over 8 perturbed inputs), inside the
1e-4 margin on the CI GPU. Seed 7 moves it by 2.3e-3 to 2.9e-3 on every
input. Tested: test_mlx_mir.py passes.

* List the re-seeded MLX MIR test in the changelog

* Refuse a zero-norm component in normalization

normalize_components now raises ValueError on a zero or non-finite
row norm instead of returning NaN. The reference path cannot reach it
(the unit diagonal keeps every norm >= 1). Tested: zero and NaN rows.

* Check a supplied NumPy A at fit start

A supplied A (or one left by a previous fit) now raises ValueError on
a wrong shape, a non-finite entry or a singular model block, through
the shared validate_supplied_mixing. Before, it raised LinAlgError deep
in the fit, or (NaN) ended with no stop reason. Tested: all three.

* Test NumPy fix_init and the wrapper round trip

fix_init (an attribute) starts from the stacked identity and leaves
the generator untouched; AMICA save/load restores the PyTorch A bit
for bit. Tested: test_initial_mixing.py passes.

* Separate the freeze from the collapse figures

The 9.0e-4 gm drop is pamica's own fit, whose A is held on those
iterations; the 4.2e-3 both sides reach comes from the freeze-free
continuation. Docs and test comment now say so. Docs only.

* Tighten three comments from the review

ADR 0006 marks the init difference as former; the _FLOOR_MARGIN
comment names the worse of doscaling on and off; pre_change notes a
component is a block row in both layouts. Comments only.

* Record the supplied-A check in the changelog

* Sphere both oracle sides with the reference S

The merged-state oracle now sphers pamica's data with the binary's own sphere, so both sides start from the same bits after the Phase 12 and 13 merge. Tested: gated component-row oracles pass.

* Start the pre-344 control from the new init

The pre-#344 control class predates #341, so it now starts from the normalized initial A the seeded reference uses. Tested: gated family oracles pass.
* Delete the dead legacy pdf code in numpy_impl/pdf.py

choose_pdf_type had no callers, and compute_pdf's pdftype 2-4 branches
were unreachable: every caller uses the generalized Gaussian. Tested
with the pdf, reference-constant and viz suites.

* Correct public docstrings against the epic code

Rewrap docstring lines that began with an issue number, which the API
pages rendered as headings; state the sphered-space shapes of the
wrapper's mixing and unmixing accessors; name restart_error among the
degenerate stops; drop the claim that natural gradient alone falls short
of the reference; fix the NumPy n_channels and doscaling wording.
Checked on the bundled sample (accessor shapes and composition).

* Say that maxdecs counts decreases, not consecutive ones

The reference resets its decrease count only at each ratchet and when
Newton switches on (amica15.f90:1062-1076), as every backend does.

* Describe the reference's iteration order in the concept pages

How AMICA works now walks one iteration as the code runs it: the E-step,
the likelihood-decrease response and stopping checks, the exit before
any update, then the update with Newton, the A-freeze windows and the
doscaling rescale. The false claim that every iteration raises the
likelihood is gone, and the flow figure shows the checks before the
update. What is AMICA covers the unit-norm scale convention.

* Fix the EEGLAB guide's scalp-map example and multi-model note

The variance-order example took scalp maps from the sphered-space
get_mixing_matrix; it now uses get_sensor_mixing_matrix. The multi-model
note still said the per-model layout differs from the reference's, which
#159 and #334 fixed. The guide index lists all four guides.

* Correct the backends guide on devices, CUDA float32 and scope

Auto device selection picks MPS first and the wrapper redirects it to
CPU for float64; the raw AMICATorchNG raises there instead (checked on
this host). CUDA float32 is about as fast as float64, not faster, as the
same page says above. The native backend is listed, the wrappers' two
backends and their parameter-file route are stated, and a stale 'sweep
is being finalized' note points to the published sweeps.

* Document the raw PyTorch backend's automatic device choice

With device=None it picks MPS first and, at the float64 default, raises
on an Apple machine; the AMICA wrapper redirects to CPU. Checked on this
host.

* Update install, accessor and precision notes on the entry pages

pamica is on PyPI (0.3.3), so getting started no longer says a release
is planned, and says unreleased changes need a source install. The
README's get_mixing_matrix comment called the sphered-space matrix
scalp maps; the examples now use get_sensor_mixing_matrix for maps. The
home page no longer calls float32 faster, the final_ll_ note covers the
max_iter case, and the NumPy CLI takes either parameter-file format.

* Correct the API overview and backend pages

List AMICANative in the public import surface; install extras with uv
from PyPI or a source checkout; correct the NumPy page's n_channels
claim (the raw PyTorch and MLX constructors take it) and list what that
backend lacks; drop the viz page's reference to the closed #159 as an
open question; replace em-dashes in list items.

* Say the reference compiles Newton off, not on

amica15_header.f90 declares do_newton = .false.; only the bundled
input.param turns it on. The Newton docstrings and the concept page
said the reference runs Newton by default.

* Correct the differences guide's Newton row and backend table

The reference compiles do_newton off (only the bundled input.param
turns it on); row 5 names singular_ll; the backend table gains the
variance_order, scoring and wrapper rows and the wrapper .pt save; an
obsolete pre-#334 comparison of two retired sharing metrics is replaced
by the current single-kernel account; the LLt exception names the last
iteration; table cells say none instead of an em-dash.

* Bring the validation guide's non-numeric claims up to date

The degenerate-fit row still said raw backends were ungated (#306 is
done) and named only the likelihood stops; the EEGLAB section still said
multi-model files are not in the reference layout. Parity figures are
left to #351.

* Make the Unreleased changelog one coherent account

Group the phase entries by what changed (fitting, MLX and the wrappers,
the MNE wrapper, persistence and exports, fixes, removals, docs) and
open with a warning that default fits differ from 0.3.3 and why. Drop
the entries a later phase superseded (Phase 8's pending #344 note, the
polish round's harness and Phase 1's remaining-gap notes), merge the
duplicate pdftype-default entries, correct the n_channels claim, and
mark measurements taken before later phases. No figure is changed.

* Map the shared modules and the epic's fitting changes in AGENTS.md

Add rank, schedule, initialization, component_layout and
reference_constants (with blocktune, restarts and the params reader) as
the shared-decision modules, plus mne_compat, native and metrics; name
amica15.f90 as the reference source; add a status line for the
reference-fidelity changes and the #306/#339 degenerate-fit extensions;
shorten the epic #278 history sentence.

* Bring the context plan and progress summary up to date

Record epic #324's phases and the reference-fidelity changes, the
#306/#339 degenerate-fit extensions and the #334 sharing oracle; close
the stale release items (repository transfer, docs deploy, PyPI and
Zenodo are done) and the #15 coverage item.

* Document the opt-in native-binary tests and MNE's pre-whitener

The testing page names AMICA_RUN_FORTRAN=1 and the mne_tests directory;
the MNE page's source formula includes the per-channel-type scaling the
export applies.

* Say shared component, not shared column, in AMICAICA's docstring

* Note that doscaling is on by default in What is AMICA

* Avoid em-dashes in the plan's updated phase headings

* Mention the post-update non-finite stop in How AMICA works

* Reword the paper's framing in nuanced terms

Drop negation-contrast and candor framing outside the Validation
section, count releases as of today (nine, six on PyPI), and describe
the MEG status as end-to-end use without measured parity.

* Drop candor framing in the differences guide and ADR 0007

* Rebuild paper.pdf

* State the complete transform in get_unmixing_matrix

The docstring called W @ get_sphere() the full transform; transform
also subtracts the mean before sphering and the model center after.
The MIR docstrings and the metrics page now call W @ sphere the linear
part. Checked on the bundled sample with two models: the full formula
matches transform exactly, W @ sphere @ X is off by up to 4.2, and MIR
is unchanged by a shift of the data.

* Say which default fits #335 and #334 reach, in the changelog

The opening warning said the schedule gates and component rows matter
only with Newton, rejection, restarts or sharing; they also reach plain
default fits in narrow cases (a maxdecs ratchet past newt_start, and
the per-component ndtmpsum sum near a grad-norm stop). The stopping
iteration's skipped steps are listed in code order (kurtosis switch,
share scan, MIR waypoint, rejection), in ADR 0008 too, and the #304
entry says the fit applies the file's settings.

* Replace the two em-dashes left in the validation guide

Prose only; no figure on those lines changed.

* Say that fit applies from_params_file's settings

from_params_file stores the file's settings on the instance; fit passes
them to the backend constructor as defaults and emits the warning.

* State effects directly in the prose this PR added

Reword contrast framings in the concept page, home page, viz page and
module docstring, and the differences guide's opening, per the style
rule: say what the code does and its known effect.

---------

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Take the wrapper's fit defaults from the backend

AMICA.fit restated max_iter/lrate/do_mean/do_sphere/do_newton, and its
lrate (0.05, runamica15.m's) differed from every backend's 0.1. It now
reads them from the selected backend's signatures, so AMICA and AMICAICA
fits without lrate run at 0.1. Tests for torch and MLX: resolved
defaults, a default wrapper fit equals the default backend fit, and one
set of constructor defaults across the backends.

* Move the MPS/float64 device fallback into AMICATorchNG

AMICATorchNG(n_channels) with default arguments raised ValueError on
Apple Silicon: device=None picked MPS, which has no float64. The
constructor now uses the CPU for that case and logs a warning; the
wrapper and AMICA.load pass device through and drop their copy. An
explicit device='mps' at float64 still raises. Torch-only (MLX has no
device choice). Tested on MPS: raw construction and fit, float32 auto,
explicit MPS, load with device=None.

* Remove the dead pamica/data package-data entry

No pamica/data directory exists. Built with uv build: the wheel's file
list is unchanged, it imports from a clean venv, and
numpy_impl/params.json is inside.

* Document pamica's defaults and the backend device choice

Add a defaults table to the differences guide: pamica against the
compiled amica15 binary and EEGLAB's runamica15.m, why pamica follows
the binary, and how to reproduce an EEGLAB run. Update the pages that
stated the wrapper's lrate or the raw backend's MPS raise. mkdocs build
--strict passes.

* Add the Phase 17 changelog entries

The wrapper lrate change joins the 'Default fits differ from 0.3.3'
warning; entries for the backend device fallback and the package-data
cleanup.

* Give AMICA_NumPy the shared default settings

params.json set do_newton on and max_iter 2000, the only shared settings
where NumPy differed from AMICATorchNG; it now sets them off and 100, as
does the constructor's max_iter fallback (reached by the CLI). Fits
shorter than newt_start are byte-identical. New test_default_settings.py
holds NumPy and MLX to every shared AMICATorchNG default (the MLX check
moves there from test_wrapper_backends.py).

* Correct the degenerate-fit test's measured stop

The comment said torch stops on nan_ll at iteration 3; both backends
stop on nan_params with iteration 2 in all 16 measured runs, at either
lrate. The test asserts only a degenerate stop.

* Document the NumPy backend's shared defaults

Defaults section, validation guide's stops table, NumPy API page, and a
changelog entry (NumPy joins the opening warning's defaults line).

* Give AMICANative pamica's shared defaults

The engine wrote the bundled input.param's values (lrate 0.05, Newton
on, max_iter 2000, block_size 512, ...). It now takes every setting the
binary has a keyword for from AMICATorchNG's signature through
FORTRAN_TO_PAMICA_KEY; binary-only knobs are unchanged. The default
block is pamica's 8192 samples split over max_threads and capped at the
data, since the binary's block_size is per thread and a block longer
than the data gives an all-NaN fit (#292). The native oracles now start
from the bundled input.param itself (byte-identical files). Tested with
the v0.3.3 binary: native, oracle and default-settings tests pass.

* Document AMICANative's shared defaults

AMICANative joins the defaults table's pamica column; the section and
the native API page cover the settings it cannot write, the per-thread
block size, and running the binary as an input.param configures it. The
changelog records the behavior change.

* Say where AMICANative's default pcakeep comes from
* Add issue-351 Newton independent-seed study script

* Add hallu driver for the Newton seed study

* Record harness, same-start and fixture parity runs

* Record share_comps counts and float32 agreement runs

* Update share_comps merge counts in differences guide

* Re-state the harness rows with same-start parity

* Add an MLX float32 option to the Newton seed study

* Mark the issue-27 and issue-51 records as pre-epic

* Update the harness figures in AGENTS and progress summary

* Add ensemble scripts and dimsweep LL records

* Add the hallu CUDA dimsweep LL record

* Re-state the cross-backend LL table with the chaos caveat

* Reword the harness and cross-backend notes

* Pin the ensemble lrate in reproduce_table1

* Add basin and pre-epic ensemble scripts

* Record the bundled tier, basin, ensemble and MLX Newton runs

* Add the equivalence-margin check on saved ensembles

* Record the seeded keep_best ensemble

* Add the re-measured keep_best figures to ADR 0003

* Name the pre-change code precisely in validation guide

* Regenerate the multi-model ensemble figure

* Rebuild paper.pdf

* Re-state the bundled single-model and multi-model results

* Update keep_best and multi-model figures in AGENTS and guides

* Update the fixture parity figures in AGENTS and progress summary

* Record the bundled tier wall-clock on the M4 Pro

* Update the cross-backend LL claim in the backends guide

* Record the Amari pairing-correction measurement

* Update paper Table 1 bundled rows and the float32 paragraph

* Clarify the benchmark excerpt in the paper

* Rebuild paper.pdf

* Record the gated oracle and full-suite runs

* Add the Phase 15 re-measurement to the changelog

* Record the Newton independent-seed fits

* Record the same-session 70-channel timing check

* Record the unseeded-start check of the issue-145 binary

* Mark the sections of the validation guide that predate the epic

* Pin every reference setting in the Table 1 and run scripts

* Re-state the Newton-enabled basin results

* Give both matching seeds in the paper

* Rebuild paper.pdf

* Document the ensemble's invsigmin difference

* Smooth the ensemble protocol sentence

* Address the docs-claims review findings

* Address the paper review findings

* Give each ensemble distribution its own line style

* Label the external single-model figures as pre-epic

* Reflow the README parity sentence

* Record the external tier run

* Fill the external-tier figures and fix rounding slips

* Rebuild paper.pdf

* Save the pinned parameter files before Phase 17

* Re-check the reference pins after Phase 17

* Pin do_approx_sphere in the reference settings

* Record the test runs after the Phase 17 merge

* Add the issue-351 findings record

* Reword two remaining contrast phrasings in the paper

* Rebuild paper.pdf

---------

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Filter CI on code paths (docs, paper, paper.pdf, metadata and
non-Python .context records skip; .context Python files and the
tested ensemble.npz still run), and commit the rebuilt paper.pdf
back only on dev and main. Checked with actionlint.
@neuromechanist
neuromechanist merged commit d79328b into dev Sep 24, 2026
10 checks passed
@neuromechanist
neuromechanist deleted the feature/issue-324-epic-mlx-first-class branch September 24, 2026 07:19
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment