Skip to content

Release 0.4.0 - #368

Merged
neuromechanist merged 69 commits into
mainfrom
dev
Sep 24, 2026
Merged

neuromechanist merged 69 commits into
mainfrom
dev

Conversation

@neuromechanist

@neuromechanist neuromechanist commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Release 0.4.0. Merge this with a regular merge commit. auto-tag.yml then strips .devN, tags v0.4.0 and creates the GitHub release, using the changelog's 0.4.0 section as the notes. That starts publish.yml (PyPI), release-binaries.yml and sync-dev.yml.

0.4.0 at a glance

MLX becomes a first-class backend, reachable through every wrapper feature (epic #324, completing the raw backend of epic #278),
and every backend's fitting follows the Fortran reference more closely.
Against the pinned v0.3.3 reference binary, single-model fits of a 70-channel EEG recording (5 seeds, 2000 iterations)
now agree with it to 6e-6 in log-likelihood, with a mean matched component correlation of 0.9996
(Validation & Parity).

The full release note is the 0.4.0 section of the changelog. It covers:

Pre-release checks

  • The version on dev is 0.4.0.dev0 (set in Prepare the 0.4.0 release #367). CITATION.cff and .zenodo.json agree.
  • The changelog carries ## 0.4.0 - 2026-09-24 and no ## Unreleased. This PR's Changelog check verifies both on the merged tree.
  • Rulesets: main has deletion,non_fast_forward,pull_request; dev has none. The AUTO_TAG_PAT secret is present.
  • dev CI is green on the epic merge (d79328b) and on the last code PRs.

neuromechanist and others added 30 commits August 25, 2026 03:13
* Add transform and mixing accessors to AMICAMLXNG

Ports AMICATorchNG.transform plus get_mixing_matrix/
get_unmixing_matrix/get_sensor_mixing_matrix/get_rho to the MLX
backend (issue #287, epic #278 Phase 1). transform derives the
unmixing composition from MLX's own _forward rather than
transcribing torch's tensor layout: MLX's W is (n_models, n, n),
not torch's (n, n, n_models). Fitting is untouched.

* Add state_dict/save/load persistence to AMICAMLXNG

state_dict/from_state_dict mirror AMICATorchNG's 12-name param set
and refusal guards (unfitted, degenerate, non-finite); save/load
write a single device- and framework-agnostic .npz (config/extra as
JSON-encoded scalars, params as native arrays -- no torch coupling,
no pickle). Loading rebuilds the MLX-only per-iteration caches
(lgamma table, log|det W|, comp_used) that AMICATorchNG instead
recomputes inline on every call.

* Add MLX transform tests

MLX-only mechanics for issue #287: model_idx validation, unfitted
errors, and transform(training X) reproducing the fit's own
E-step activations via the transpose identity _forward relies on.
Includes the fit-path no-op pin (C7): a short seeded fit is
bit-identical to a recorded run from the epic branch tip, proving
this phase touched nothing _fit_once calls.

* Add MLX transform cross-backend tests

Pins transform and the get_mixing_matrix/get_unmixing_matrix/
get_sensor_mixing_matrix/get_rho accessors against a float64
AMICATorchNG twin holding identical fitted parameters, following
the existing test_mlx_newton_cross_backend.py pattern. Covers
multi-model model_idx routing with nonzero c, and a genuinely
rank-reduced fit for get_sensor_mixing_matrix.

* Add MLX persistence tests

Round trip (state_dict/from_state_dict and .npz save/load): every
param bit-identical, config/extra equal, transform output
bit-identical pre/post. Sharing and adaptive-switcher state round
trip (forced merge, pdftype=1). Refusal guards: unfitted,
degenerate, wrong format_version, missing sections, shape drift.

* Document MLX transform and persistence support

Update the AMICAMLXNG module docstring, docs/api/mlx-backend.md,
docs/guides/amica-differences.md's backend table, AGENTS.md's MLX
status sentence and the changelog to reflect issue #287 (epic #278
Phase 1): transform/accessors/save-load are no longer "fast-follow"
gaps. Remaining gaps are keep_best (Phase 2) and outlier rejection
+ LLt/MIR (Phase 3).

* Persist n_newton_fallbacks in AMICAMLXNG state_dict

Plan enumeration for extra omitted this field (an oversight, not a
deliberate divergence -- torch's equivalent extra dict includes
it); a reloaded model was silently losing the Newton-fallback
counter. Round-trip test drives a genuine nonzero count via a
deliberately under-determined 256-sample block (same construction
as test_mlx_newton.py::test_fallback_ramps_toward_lrate_cap_and_counts)
through both state_dict/from_state_dict and .npz save/load.

* Harden AMICAMLXNG persistence load path

PR review on #305 found the load path under-validated:
- _load_params shape-checked only A/comp_list; extend to all 12
  _PARAM_ARRAYS against an expectation table read off
  _initialize_parameters's real allocations. sphere's input-channel
  width isn't derivable from config under rank reduction, so it's
  taken from the restored sphere itself and mean is cross-checked
  against it.
- comp_list/pdtype no longer force-cast unconditionally: a float
  array with NaN or a fractional value is numpy UB once cast to an
  index dtype. _safe_int_cast requires the value already be integer,
  or finite and whole, before casting; corrects the "dtype
  preserved" docstring claim (it force-casts).
- load() wraps the npz read so a byte-truncated or otherwise corrupt
  file (zipfile.BadZipFile/EOFError/OSError/ValueError) re-raises as
  the same named "malformed AMICAMLXNG save file" error instead of
  an opaque exception.
- extra now requires all 14 keys (format_version 1 has no older
  files to be lenient for), mirroring the params completeness check,
  instead of 10 keys raising bare KeyError and 3 silently
  defaulting to [].
- W recomputes log|det W| on load; a finite-but-singular restored W
  makes slogdet return -inf rather than raise, so that's checked
  explicitly.
- Dropped the dead _INT_PARAM_ARRAYS tuple (only _INT_PARAM_DTYPES
  is read).

Also adds AMICAMLXNG.n_channels_in (port of AMICATorchNG's property),
so a reloaded rank-reduced model can still report its original input
width, and corrects the stale "this backend has no load path" comment
on _sphere_np now that _load_params is a second sphere-assignment
point.

* Fix stale MLX backend-parity example in rules doc

.rules/backend_parity.md's "existing legitimate examples" sentence
still said transform raised NotImplementedError and save/load was
absent; both landed in epic #278 Phase 1 (#287). Names only the
genuinely still-absent items: do_reject, keep_best, LLt/MIR
(Phases 2/3).

* Add persistence hardening and review-requested test coverage

Covers the PR #305 review fixes:
- shape-drift rejection for every param (not just A/comp_list),
  including mean's sphere-derived n_channels_in cross-check
- comp_list dtype safety: rejects NaN and fractional values, accepts
  whole-valued float arrays
- singular restored W is rejected (a whole zeroed column reliably
  reaches slogdet's -inf; a duplicated column on an already
  float32-rounded W does not -- documented in the test)
- a real byte-truncated (not just well-formed-npz-missing-a-key)
  .npz raises the same named error
- get_rho() and state_dict() defense-in-depth: a force-set
  non-finite rho/param (clean stop_reason) is caught and named
- n_channels_in round trips on a rank-reduced fit
- rank-reduced save/load round trip: all 12 params + transform
  bit-identical through both state_dict and .npz
- n_restarts=3 round trip: restart_seeds_/restart_lls_/
  restart_stop_reasons_ keep their full per-restart length/content,
  not a vacuous length-1 list
- cross-backend accessor test: documents why get_unmixing_matrix
  alone needs the looser 1e-3 bound (independent per-backend matrix
  inversion, not a pure reindex/transpose like mixing/rho)

* Make fit-path no-op pin survive cross-GPU float32 noise

CI failed on a different Apple GPU model: MLX float32 is
bit-reproducible on one machine but not across GPU models
(observed: ll_history[0] differs from the M4 Pro recording by
~7e-8 relative). Reshape the pin instead of deleting it:

- ll_history/final_ll_ now compared per-entry within 5e-6 relative
  (~50x the observed ~1e-7 cross-GPU spread, below the ~1e-6+ shift
  a genuine fit-path change produces -- see #216's block_size
  precedent) instead of exact equality.
- Drop the SHA-256 A/W hashes (a cross-machine hash can never
  match); replace with the same 5e-6 tolerance on five
  representative A entries.
- stop_reason and len(ll_history) stay exact -- machine-independent.
- Docstring now says plainly that the bit-level no-op claim was
  verified same-machine at development time (epic tip 076605c,
  before/after, bit-identical there); this test is the looser
  cross-machine canary.

* Scale the A-entry no-op spot checks by matrix max, not per-entry

CI failed again: A[10, 20] is small (-0.0037), so the per-entry
5e-6 relative bound from the previous fix was an absolute tolerance
of ~1.9e-8 there -- tighter than the ~1.1e-7 cross-GPU float32
noise CI actually observed. Same near-zero-entry failure mode
_max_rel_disagreement's docstring in
test_mlx_transform_cross_backend.py already explains for
transform's output; apply the same fix here. The five A spot
entries are now compared against a recorded max(|A|) (~1.0, A
columns are ~unit-normalized) instead of against each entry's own
magnitude. ll_history/final_ll_ keep the per-entry relative check
(safe: their magnitude, ~3.3, stays well clear of zero).
test_run_fortran_rejects_sub_resolution_config bet that 8ch x 2000
samples always stamps 0.00 s per iteration; a loaded weekly-run
macOS runner crossed one 0.01 stamp and the guard rightly had
nothing to reject (issue #308, first weekly failure). The test now
asserts the guard's actual contract -- either the named
sub-resolution error, or a mean at least one stamp over the timed
iterations, never a bogus 0.0 -- so both legitimate outcomes pass.
Verified against the real bundled amica15mac binary locally.
* Add keep_best best-iterate safeguard to MLX backend

Ports #51's snapshot/restore mechanism from AMICATorchNG onto
AMICAMLXNG: fit() tracks the highest-LL iterate and restores it when
the run ends more than _KEEP_BEST_TOL below that peak. Inactive under
share_comps (a merge changes the parameter count, so an earlier
snapshot cannot be reverted to without undoing it). Persisted
additively in state_dict's config; ll_history is never rewritten.

Epic #278 Phase 2, issue #288.

* Add keep_best tests for MLX backend

Covers the monotone no-op pin, a genuine forced-restore overshoot
(measured, not assumed), ll_history immutability, bit-identical
truncated-refit comparison, n_kurt_done/pdtype consistency under
pdftype=1, the share_comps disable, and state_dict persistence
(including additive-compat for a payload missing the key). All
bit-identity checks run within a single test process to stay safe
against MLX's cross-GPU float32 non-reproducibility.

A cross-backend companion asserts the shared restore CONTRACT
(final_ll_ == max(ll_history) when active) without asserting decision
equality between PyTorch's float64 and MLX's float32 trajectories.

Epic #278 Phase 2, issue #288.

* Document MLX keep_best in the backend-parity docs

Backend table, gaps sentences (AGENTS.md, mlx-backend.md,
backend_parity.md) and changelog now reflect keep_best as
implemented on MLX; the remaining MLX gap is do_reject + LLt/MIR
(Phase 3).

Epic #278 Phase 2, issue #288.

* Test the degenerate-stop guard and restart record

PR #310 review additions: the restore guard's refusal to rescue a
diverged fit that peaked earlier is now pinned in both the MLX and
torch backends via the sanctioned error-injection subclass (real
iterations, then a forced nan_ll, asserting no restore and bit-
identity with keep_best=False); and the keep_best x n_restarts
composition is pinned (restart_lls_ reports the restored iterate,
bit-identical to a standalone fit's final_ll_). The grad_norm-stop
recipe variant is deliberately skipped: the guard is a set-membership
test on stop_reason, so nan_ll alone covers the branch. Tested:
13 keepbest tests + torch twin pass; ruff and ty clean.
* Add LLt stash to the MLX backend

Port issue #157's per-block E-step stash (AMICATorchNG) so the LLt
export never needs a second forward pass. _get_block_updates emits
logV/ll_samples per block, _accumulate_blocks(stash_llt=True)
scatters them into _llt_logv/_llt_ll as part of the same lazy graph,
and _fit_once materializes them into _llt_lht/_llt_lt after any
keep_best restore (which now rolls the stash back with the params).
Fixes three test doubles that override _accumulate_blocks to accept
the new stash_llt kwarg.

* Add do_reject outlier rejection to the MLX backend

Port issue #123's AMICATorchNG good_idx mechanism: constructor gains
do_reject/rejsig/rejstart/rejint/maxrej with torch's names, defaults
and validation; _fit_once shrinks good_idx on the reject schedule and
restricts the E-step to it; keep_best's track_best now also excludes
do_reject.

Design decision: the rejection statistic is read FROM the LLt stash
(_llt_ll indexed by good_idx) instead of a second forward pass, the
NumPy backend's design -- pre-empting AMICATorchNG's open follow-up
(#298) to drop its own extra _sample_ll pass. numrej/good_idx persist
additively in state_dict's extra (extra.get fallback), the one place
phase-1's require-all-extra-keys check is relaxed, documented at both
call sites.

* Add model_loglik/model_probability to the MLX backend

Port issue #141's per-model posterior accessors (AMICATorchNG
core.py:3041-3132): score arbitrary data through the stored
sphere/mean/W via the existing block-wise _forward, returning
(n_models, n_samples) Lht, plus the column-softmax model_probability.

* Add write_amica_output EEGLAB export to the MLX backend

Thin adapter over numpy_impl.load.write_amicaout, matching
AMICATorchNG.write_amica_output (issue #92): converts mx arrays to
numpy, transposes W from MLX's (n_models, n, n) to write_amicaout's
(nw, nw, num_models) contract, applies the keep_best LL-truncation so
the written LL trajectory ends at the returned iterate, and passes
Lht/Lt from the materialized LLt stash (warn-and-omit for a
from_state_dict/load model). Verified against a float64 torch twin
built from identical fitted parameters: W/A/LL/LLt/comp_list agree to
float32 precision (~3e-7) or exactly.

* Add mir/pmi diagnostics to the MLX backend

Port issue #137's MIR/PMI accessors (AMICATorchNG core.py:2929-3036)
delegating to pamica.metrics: mir() composes
get_unmixing_matrix(model_idx) @ sphere, pmi() runs pairwise_mi on
transform(). Ports the #300 fitted-geometry PCA guard
(sphere.shape[0] != sphere.shape[1]) rather than a config-only check,
since this backend has no explicit pcakeep/pcadb parameter. fit(X,
mir_step=...) appends (iteration, mir_nats, variance) waypoints to
mir_history_ on the given schedule, excluded from the keep_best
snapshot and from state_dict (a true trajectory, not a fitted
parameter).

* Add tests for MLX reject/LLt/export/scoring/mir/no-op

Six new mlx_tests files covering the phase's five deliverables plus
the fit-path no-op check, and the one new cross-backend file
.rules/backend_parity.md calls for:

- test_mlx_llt_stash.py: the Lt.sum()/(n_good*nw) == final_ll_
  invariant, keep_best stash rollback, LLt omission with no E-step.
- test_mlx_reject.py: good_idx schedule, maxrej cap, none/all-rejected
  edges, keep_best inactive warning, state_dict/save-load round trip
  and additive-fallback loading of pre-Phase-3 payloads.
- test_mlx_export.py: write_amica_output -> loadmodout round trip,
  do_reject zero-sentinel, from_state_dict LLt omission, no extra
  forward pass, and a byte-level comparison against a float64 torch
  twin's export.
- test_mlx_scoring.py: model_loglik/model_probability against the
  E-step's own forward pass and the one-M-step-behind stash.
- test_mlx_mir.py: mir/pmi composition, mir_step schedule, the failed-
  waypoint NaN guard, the #300 PCA-reduction guard on a real
  rank-reduced fit, and mir_history_'s keep_best/state_dict exclusions.
- test_mlx_fit_noop.py: a default fit (do_reject off, mir_step 0) is
  bit-identical, same process, to the epic tip before this phase
  (loaded live via git show as a sibling module).
- test_mlx_reject_cross_backend.py: MLX and PyTorch reject the SAME
  sample set from identical config (evidence for the stash-based
  rejection-statistic design decision), plus mir() numeric agreement
  against a torch twin.

* Document epic #278 Phase 3: MLX gaps closed

Update the MLX backend module/package docstrings, docs/api/mlx-backend.md,
AGENTS.md, docs/guides/amica-differences.md (backend table, do_reject/LLt/
MIR sections, a new #298 rejection-statistic design-decision section) and
.rules/backend_parity.md to reflect that outlier rejection, the LLt stash,
the EEGLAB export, and MIR/PMI are now implemented on MLX, closing epic
#278. Adds the Unreleased changelog entry.

* Fix test_mlx_fit_noop.py for shallow CI checkouts

git show 2e04006 fails with "bad object" in CI's depth-1 checkout,
erroring all three no-op tests. Fetch the one commit needed first
(tolerating failure), and skip loudly naming the shallow-clone cause
if the object is still unreachable, instead of letting git show raise.

* Fix _choose_pdfs to use the do_reject good set on MLX

_choose_pdfs(X_t) passed the full sphered dataset instead of X_use
(the do_reject-restricted good set AMICATorchNG passes), so under
pdftype=1 + do_reject the kurtosis-based family decision silently saw
the outliers do_reject had already excluded. Fixed to pass X_use.

Regression test in test_mlx_reject_cross_backend.py: a pdftype=1,
do_reject=True fit's per-source pdtype decision now matches a float64
torch twin bit-for-bit (confirmed to mismatch 1/32 before the fix, 0/32
after).

* Guard write_amica_output against degenerate/non-finite models

write_amica_output had no refusal guard on either backend: a degenerate
(non-finite-LL) or force-corrupted model could be written silently.
state_dict() already refuses via a two-layer guard (degenerate
stop_reason, then a defense-in-depth isfinite sweep over the param
arrays); add the identical guard to AMICATorchNG.write_amica_output
and AMICAMLXNG.write_amica_output. The AMICA wrapper's usability gate
only protects callers going through it, so a caller using either
backend class directly had no gate at all -- and neither backend has
one now missing.

This also neutralizes the stale-NaN-stash path: a failed final
iteration's LLt stash (from before a nan_params/nan_ll break) can no
longer reach disk, since the write is refused before the stash is
ever read.

Tests: direct-marker degenerate refusal + force-corrupted-param
defense-in-depth, one pair per backend (torch: test_ng_backend.py;
MLX: test_mlx_export.py, NaN-injection recipe matching
test_mlx_persistence.py's existing state_dict test).

* Pin do_reject through best-of-N restarts and share_comps

Two interaction regression tests (PR #311 review), both verified to
hold before being pinned:

- do_reject x n_restarts: with the winning seed forced to NOT be the
  last restart (SEEDS=[54,50,42], test_mlx_restarts.py's existing
  winner-first fixture), good_idx/numrej/A match a solo fit from the
  winning seed bit-for-bit -- the restore path (_apply_restart_state)
  is what a last-restart winner would leave untested.
- do_reject x share_comps: comp_thresh=0.9 (test_mlx_sharing.py's
  real-data recipe for a genuine, non-forced merge) combined with a
  do_reject schedule that also fires in the same fit. Asserts finite
  params, the merged comp_list/comp_used surviving, and good_idx
  shrunk alongside it.

* Fix mir_step docstring and document the auto-rank-reduction parity

PR #311 review: no behavior change. The mir_step docstring's "cannot
be known before _preprocess" framing was misleading -- the sphere's
shape IS knowable right after _preprocess runs, a few lines later. The
real reason there is no upfront gate for automatic mineig/mineig_rel
reduction is torch-parity: AMICATorchNG's own upfront gate only ever
covers an explicit pcakeep/pcadb request (issue #300's deliberate
scope), never the automatic case, which both backends only catch
downstream, per-waypoint, inside mir()'s own #300 geometry guard.
Reworded the docstring and added a clarifying section to
amica-differences.md.

Adds the missing test: fit(mir_step=N) on rank-reduced real data (SVD
projection) completes normally, with every waypoint recorded as a
warned (iter, NaN, NaN) entry.

* Match torch's keep_best-inactive reason precedence exactly

MLX checked share_comps before do_reject when reporting why keep_best
is inactive; AMICATorchNG checks do_reject first
(torch_impl/core.py:2425). When both flags are on the two backends
must report the same reason for the same configuration. Swapped to
match, and pinned with a test where both are enabled simultaneously.

* Extend the torch-twin export byte comparison to n_models=2

PR #311 review: parametrize the single-model-only byte comparison to
also cover n_models=2, and add an explicit per-model W diff (reshape
the raw file back to write_amicaout's own (nw, nw, num_models)
contract and compare model 0 and model 1 slices separately) so a
per-model swap or a wrong axis in the (n_models, n, n) -> (n, n,
num_models) transpose shows up as a named model index, not just
"W differs somewhere" in the aggregate flat-byte check. Confirmed
non-vacuous: the two models' W differ by ~0.16 in this fit.

* Document the direct-call injection-pattern variant in testing.md

PR #311 review: test_mlx_reject.py and test_numpy_reject.py already
use a lighter variant of the sanctioned error-injection pattern --
calling _reject_outliers directly on a hand-built ll_vec, rather than
a subclass, to reach the all-rejected branch a real fit can never
organically trigger. Acknowledge it explicitly alongside the
subclass form.

* Mention the write_amica_output guard in the changelog

Amends the Phase 3 Unreleased entry to note the degenerate/non-finite
refusal guard added to write_amica_output on both backends (PR #311
review, item 3).
* Attribute degenerate-fit refusal to AMICA wrapper

Row 5 of the differences table credited the raw PyTorch backend with
refusing transform/get_*/save on a degenerate fit; that guard is the
AMICA wrapper's _check_usable contract. The raw AMICATorchNG backend
has no such guard of its own (tracked as issue #306).

* Document unmapped Fortran keywords

Add a section resolving the page's exhaustiveness contract: points to
fortran_params.py's FORTRAN_UNSUPPORTED_KEYS as the enumeration, names
filter_length/dft_length/decwindow as dead in the reference itself
(alongside the already-documented do_choose_pdfs/Spinv2 dead code),
and records the do_rho-vs-pdftype divergence: Fortran can freeze the
GG shape while keeping pdftype=0, but every pamica backend derives
the freeze from pdftype directly, making that combination unreachable.
Links issues #312/#314 for the live lifecycle/diagnostics gaps.

* Port variance_order to AMICAMLXNG

Adds the EEGLAB back-projected-variance component order (issue #92)
to the MLX backend, mirroring AMICATorchNG.variance_order with MLX's
model-major W layout ((n_models, n, n)) and the unfitted RuntimeError.
Closes the one accessor gap the epic's feature-parity audit found in
an otherwise-complete Phase 3; updates the module docstring's
ported-features claim to match.

Adds cross-backend parity coverage against a float64 AMICATorchNG
twin holding identical fitted parameters (order matches exactly on a
30-iteration multi-model fit with non-degenerate variance gaps, with
a gap-size assertion documenting the tie risk), plus MLX-only
unfitted/invalid-model_idx/permutation tests.

* Correct validation harness backend coverage claim

AGENTS.md claimed the harness runs "both implementations (NG +
NumPy)"; validate_implementations.py only exercises torch vs Fortran.
NumPy parity lives in pytest (test_sample_data_numpy_vs_fortran); MLX
validation lives in mlx_tests/ plus the cross-backend suites.
Extending the harness itself is tracked as issue #315.

* Fix unsupported-key comment for writestep/do_history

The FORTRAN_UNSUPPORTED_KEYS entries for writestep/do_history/
histstep said these have no pamica equivalent; false for the legacy
NumPy backend, which implements all three (numpy_impl/core.py's fit
loop and _write_history). Keeps the keys unsupported here (this
module targets the torch wrapper, which has no matching mechanism)
but corrects the reason: NumPy-only, torch/MLX gap tracked as #312.

* Refresh stale feature-parity status notes

feature_parity.md and progress_summary.md still claimed the runtime-
vs-Fortran benchmark was unmeasured and issue #15 (save/load and
plot_components test coverage) was open; both landed since. Also
corrects progress_summary.md's validation-harness claim (same fix as
AGENTS.md: torch-vs-Fortran only, not NG + NumPy). Adds a header note
to feature_parity.md pointing to amica-differences.md's table as the
live, maintained backend-parity source of truth.

* Align NumPy backend's inert pdftype default to 0

self.pdftype is never read after assignment (this backend always
runs the GG update; _compute_log_pdf takes no pdftype param), so the
old default of 1 was a harmless but confusing mismatch against
torch/MLX's constructor default of 0. Aligned for surface consistency
only; no behavior change.

* Add changelog entry for epic polish round

Summarizes the audit-driven fixes ahead of merge to dev: MLX
variance_order, the doc corrections, and the new tracking issues
(#312, #314, #315) plus the #306 extension.

* Scope variance_order precision claim to defaults
A RuntimeError from inside _fit_once's iteration body (the #274
condition-number guard, _choose_pdfs's pdtype invariant, and
_pinv_sphere's non-finite-sphere guard on MLX; _update_unmixing_
matrices's torch.linalg.inv failure and _pinv_sphere's guard on
torch) fires AFTER that iteration's _update_parameters already
reassigned A/mu/beta/rho/alpha/gm/c, but BEFORE W/_logdet_W are
updated. stop_reason stayed "max_iter", so a caller of the
single-restart fit() path (no try/except there) that catches the
propagated exception and persists the model anyway would have every
state_dict()/write_amica_output() degenerate-fit guard accept it
silently.

Set self.stop_reason = restarts.ERROR_STOP_REASON immediately before
each raise, so the instance itself -- not just the exception -- is
left degenerate regardless of whether the caller catches it. The
multi-restart path's own except block already did this; kept as
redundant-but-harmless so it does not depend on every raise site
doing this correctly.

Tests (both backends): a sanctioned injection subclass runs a real
fit for N genuine iterations, then forces the target guard's raise
on a later iteration -- confirmed to fail without the fix (stop_reason
stayed "max_iter") and pass with it, and that state_dict/
write_amica_output then refuse the resulting instance. Plus
direct-call coverage for the two guards not reachable via a real
fit trajectory.
_snapshot_params/_restore_params rolled back W/rho/etc via
_PARAM_ARRAYS, but the two MLX-only per-iteration caches
(_logdet_W, feeding log|det W| into every logV; _lgamma_table,
feeding the GG log-density) were never captured -- so a restore left
them at the LAST (discarded) iterate's values, silently corrupting
model_loglik/model_probability/mir on the "restored" model even
though W/rho themselves rolled back correctly.

Added both to the snapshot dict (captured unconditionally, since
both always exist once _initialize_parameters has run) rather than
rebuilding them post-restore the way _load_params does: it keeps the
snapshot a single self-contained point-in-time capture with nothing
a future _restore_params caller could forget, and _restore_params's
existing generic setattr loop handles them with no extra code at the
call site. Both are already in _RESTART_STATE_ATTRS.

Regression test (forced-restore recipe, oracle = a state_dict round
trip, which independently rebuilds both caches from the restored
W/rho): measured max_diff = 0.3804 (100% of 8192 elements, up to
~0.43% relative error) before the fix, 0.0 (bit-exact) after.
mir()'s ValueError (PCA reduction, or metrics.mir's own near-singular
-unmixing check) reflects the fit's geometry -- a fact that does not
spontaneously resolve mid-fit -- so every remaining scheduled
waypoint re-logged the identical warning on a long fit. LinAlgError
stays per-waypoint: that is the genuinely transient case (a
near-singular W the natural gradient can pass through) the existing
comment already describes.

After the first ValueError: warn once (now naming that waypoints are
disabled), record that one NaN entry, and stop scheduling waypoints
for the rest of this fit (no more mir() calls, no more entries) --
both backends. A local flag inside _fit_once, not an instance
attribute: it only matters within one fit() call.

Updated the two existing tests this changes the contract of
(injected-failure and organic-rank-reduction) to assert the new
stop-after-first-ValueError behavior on both backends, including that
mir()/flaky_mir is never called again after the failure.
_load_params's finiteness check only covered comp_list/pdtype (via
_safe_int_cast) -- the other 10 float _PARAM_ARRAYS (A, W, c, mu,
alpha, beta, rho, gm, mean, sphere) had no isfinite check at all, so a
NaN/inf-poisoned payload would load "successfully" and only surface
later as a confusing downstream NaN with no diagnostic, contradicting
the same promise the existing W-singularity check already makes for
W specifically. Extends the named-ValueError check to every float
param.

Test: a NaN-poisoned mu payload (representative of the 10) is
rejected by from_state_dict with a named error.
max_iter=0 used to run the EM loop zero times and "complete"
successfully with stop_reason="max_iter" (not a degenerate marker)
and final_ll_=NaN -- an untrained model that every state_dict()/
write_amica_output() degenerate-fit guard then accepted, since none
of them check "did an E-step ever actually run", only "did
stop_reason end up degenerate".

Added a named ValueError at fit entry on both backends, alongside
the existing X.ndim/mir_step checks -- both are argument-validation
guards raising before any work happens, and the multi-restart path's
narrow "except RuntimeError" already lets ValueError propagate on
the first restart rather than swallowing it as a degenerate one.

The NumPy backend already validates max_iter >= 1 at construction
(pre-existing, unmodified); no fix needed there.

Tests: max_iter=0 raises the named error and leaves the model
untouched (unfitted) on both backends.
_load_params's numrej/good_idx restoration reached raw int()/
np.asarray() calls with no try/except -- a corrupted or hand-edited
payload's malformed value there raised an opaque bare
TypeError/ValueError instead of the "malformed AMICAMLXNG state: ..."
message every other field in this method already uses.

Tests: a non-numeric numrej and a good_idx that cannot be converted
to an integer index array each raise the named error.
Mirrors test_mlx_reject.py's do_reject analog. mir_history_ is in
_RESTART_STATE_ATTRS, so this exercises the same snapshot/restore path
with mir_step set so waypoints actually populate (PR #318 review item 7).
__init__ already validated max_iter, but it is a plain public attribute
a caller can reassign afterward, so fit() silently ran the EM loop zero
times if max_iter was mutated post-construction. torch/mlx take
max_iter as a fit() parameter re-validated every call; numpy takes it
only at construction, so re-check self.max_iter at fit entry to close
the same gap (PR #318 review item 5 clarification: all three backends).
One fit combines n_models=2, share_comps (forced genuine merge),
do_reject, n_restarts=2, and pdftype=1 (with keep_best=True passed to
confirm the do_reject/share_comps exclusion is logged, not silently
ignored), then walks every accessor, state_dict/from_state_dict,
.npz save/load, and write_amica_output/loadmodout, asserting
consistency at each stage (PR #318 review item 8).
Both use the deterministic two-model forcing recipe from
test_mlx_sharing.py (share_start=4, share_iter=8, comp_thresh=0.9): a
genuine merge, then write_amica_output/loadmodout round trips, and
mir() stays finite and per-model_idx-aware on the merged model
(PR #318 review item 9).
_PARAM_NAMES was missing mean/sphere/pdtype, so the module docstring's
'every fitted array is compared bit for bit' claim was not literally
true. All three agree bit-for-bit with the pre-Phase-3 epic tip
(PR #318 review item 10).
Both backends' no-E-step LLt tests called fit(max_iter=0) to reach a
'never fitted' state, which item 5's max_iter>=1 guard now rejects
before that state is reachable. Rewritten to check the same
construction-time invariant directly; the write_amica_output-omits-LLt
behavior they also exercised is already covered by
test_from_state_dict_write_amica_output_omits_llt (torch) and
test_from_state_dict_model_warns_and_omits_llt (mlx).
Four places claimed 'epic #278 is complete as of Phase 3' without
naming variance_order, which actually landed in the post-Phase-3
polish round (mlx_impl/__init__.py, .rules/backend_parity.md,
AGENTS.md); also fixed docs/changelog.md's do_reject/keep_best
exclusion note from future to past tense, now that do_reject shipped
in Phase 3 (PR #318 review items 11-12).
Several docstring/comment citations into torch_impl/core.py had drifted
from their targets (this epic's own accumulated edits, plus this
round's exception-consistency/max_iter changes): write_amica_output,
_PARAM_TENSORS, the digamma dorho-gate (also pointed at the wrong
code entirely -- the NaN-reset path, not the gate itself), and the
phase-1 accessor citations (transform, get_mixing_matrix,
n_channels_in, get_rho, get_sensor_mixing_matrix, get_unmixing_matrix,
_check_model_idx). All recomputed against the current file
(PR #318 review item 13).
docs/api/mlx-backend.md listed the Phase 1 accessors but omitted
variance_order, which landed in the post-Phase-3 polish round
(PR #318 review item 14).
docs/guides/validation.md's table said rejection was 'both backends'
(stale since MLX shipped it in Phase 3) and attributed the
degenerate-fit refusal on transform/get_*/save to the raw backends,
when it is actually the AMICA wrapper's contract (the raw-backend gap
is issue #306); write_amica_output is the one exception, which gained
its own raw-backend guard directly in this epic (PR #318 review item 15).
Adds the third variant to the sanctioned exception-injection section:
wrapping a real method with monkeypatch to count/record calls while
still delegating to the real implementation, used by
test_write_amica_output_makes_no_extra_forward_pass to prove zero
E-step passes (PR #318 review item 16).
One-line addendum: the safeguard was ported to the MLX backend in
epic #278 Phase 2 (issue #288), with the same default, tolerance, and
share_comps/do_reject exclusions (PR #318 review item 17).
* Surface MLX feature parity in user-facing docs

Post-epic-#278 discoverability pass: the AGENTS.md map annotation
still said GG-only; the EEGLAB guide never mentioned the MLX export
path; the backends table row and README undersold the now-complete
MLX surface. Docs only.

* Address docs review: precision claims and naming

The export-validation claim now states float32 precision rather than
byte comparison (only comp_list is exact), and the test that seeded
the overclaim is renamed to match its actual assertion; the backends
table row returns to terse form with the parity statement moved to
prose that names the wrapper-only from_params_file caveat (#313);
MIR is spelled out on first use in the README; the stale GG-only
class one-liner in mlx core is corrected. Export suite 12 passed.
neuromechanist and others added 29 commits September 22, 2026 19:55
* feat: validate every backend in the harness

validate_implementations.py gains --backend {torch,numpy,mlx}, a comma-separated list, or all. Each backend runs the bundled sample with the harness's canonical params (NumPy through its own key map, MLX through AMICAMLXNG) against one shared reference run; an explicit --backend adds a parity summary table with runtime. The default torch-only run prints the same report as before (diffed).

Tested: dispatch and argument-parsing tests (no binary), and the AMICA_RUN_FORTRAN-gated per-backend parity test against the v0.3.3 native binary (3 passed).

* refactor: run the harness MLX row via AMICA

With Phase 4 merged, the torch and MLX runs share one AMICA(backend=...) path, so the MLX row exercises the wrapper a user calls. Numbers unchanged (same summary table as the direct AMICAMLXNG run).

Tested: harness dispatch tests (29 passed); all-backend run against the v0.3.3 native binary.

* docs: use American English spellings

35 British spellings in comments, docstrings and Markdown (behaviour, colour, neighbour, cancelled, licence, optimiser, ...). Proper nouns (Centre de Recherche...) and 'analogue' left as is; no identifiers touched.

Tested: ruff, typos.

* docs: cite torch core by symbol, not line

Seven torch_impl/core.py:<line> citations outside the MLX module pointed at drifted lines; they now name the method (or are dropped where the sentence already names it). Also replaces a stale pamica.py:802-813 citation (a file renamed in #34) in AMICATorchNG._newton_direction. Fortran line citations unchanged.

Tested: ruff.

* docs: MLX getting-started path and harness rows

getting-started gains a complete Apple Silicon (MLX) section (install, fit/transform/save/load, params file, AMICAICA, precision). The validation guide documents running the harness per backend and records the per-backend parity rows with their expected bars. The differences page records do_sphere=False (#328) and NumPy refit warm-start (#312), both verified against the code. AGENTS.md, changelog and the native-backend and testing pages follow.

Tested: mkdocs build --strict (scratch env); anchors checked in the built site.

* test: keep suite output out of the repository

An audit of the non-slow suite found 63 tests (13 files) writing ./output via AMICA_NumPy's default outdir; the slow CLI, end-to-end and pdf-family tests wrote test_output, _ng_e2e_tmp and scratch_amica_parity into the repo. A session-scoped autouse fixture now runs each session (per xdist worker) from a temporary directory, the harness resolves its bundled inputs from its own location, and the slow tests write under tmp_path. Stale .gitignore entries removed.

Tested: full non-slow suite (1335 passed, nothing written to the repo); coverage reports still land in the repo root with and without xdist.

* test: end-to-end workflow on torch and MLX

New mne_tests/test_end_to_end_workflow.py mirrors our own pipeline on the bundled EEG: average reference, AMICAICA with pcakeep = n_channels - 1 on each backend, get_sources vs transform, a single-component exclusion (exact back-projection, residual kept, rank drops by one), EEGLAB export reloaded by loadmodout with the padded sphere, AMICA save/load, an input.param-driven fit, and torch vs MLX Hungarian-matched sources >= 0.999. Not marked slow (about 5 s); MLX tests skip without MLX or an Apple GPU.

Tested: 12 passed locally (torch and MLX); measured tolerances recorded in the module.

* fix: transpose NumPy sensor mixing matrix

AMICA_NumPy.get_sensor_mixing_matrix returned pinv(sphere) @ stored A without the transpose the torch and MLX backends apply (the stored W is the unmixing transposed, issue #24), so its columns were rows of the true mixing matrix: about 10% off the torch maps after five iterations and no inverse of get_weights() @ sphere. New cross-backend test pins torch/NumPy agreement (4.6e-12 full rank, 9.5e-11 at pcakeep=20), MLX agreement at float32 tolerance, and the identity on every backend.

Tested: new test fails without the fix (4 NumPy cases) and passes with it (10 passed); backend-guard and share-comps suites green.

* chore: remove stale 2025 debug output

pytorch_debug_test/out.txt and pytorch_output_test/out.txt were run logs committed in August 2025 by the long-removed basic PyTorch backend; nothing reads them. The only reference, a typos exclusion for the first directory, goes with them.

Tested: typos.

* docs: pin the harness rows to the v0.3.3 binary

States why the rows name an explicit release binary instead of the resolver default (its cached 'latest' is never refreshed) and how to fetch it.

Tested: mkdocs build --strict.

* Rebuild paper.pdf

* fix: NumPy writes files only with an outdir

AMICA_NumPy defaulted outdir to ./output, so every fit wrote out.txt at construction, writestep checkpoints, history and its final results into the caller's working directory. The default is now outdir=None, which writes nothing (as torch and MLX never do unless asked); an explicit outdir (keyword, params file, or the CLI's --outdir, whose own default stays output) writes exactly as before. Behavior change recorded in the changelog and the NumPy API page; the conftest chdir stays as defense in depth.

Tested: new test_numpy_outdir.py (default leaves the working directory empty with every write path on; explicit and params-file outdir write as before; fails on the old code) and a CLI test that a run without --outdir still writes ./output; NumPy suites (107 passed).

* fix: NumPy viz plots the true maps and sources

plot_components drew stored-A columns (rows of the true mixing matrix) as mixing vectors, and it and plot_pdf_fits built activations as stored W times raw data (no transpose, sphere or mean); plot_model_comparison skipped the sphere. One helper, _component_maps_and_sources, now computes exactly what get_sensor_mixing_matrix and transform return, and all three plots use it. load_results reads a rank-reduced fit's zero-padded sphere instead of raising.

Tested: new test_numpy_viz.py compares the helper and the plotted lines/histograms with the live accessors, full rank and pcakeep=20 (10 passed; 8 fail on the old code).

* test: pin the default harness run and summary

Pins the literal default run (no --backend: one validation_report.txt, no parity_summary.md, exit 0), renders format_parity_summary with populated rows from two real runs (torch as the stand-in reference, NumPy through compare_results) and checks every numeric column, and runs the no-MLX exit over both 'torch,mlx' and 'all'.

Tested: 32 passed, 3 skipped (the AMICA_RUN_FORTRAN-gated parity test).

* docs: fix two more British spellings

unlabelled and relabelled, missed by the first sweep's word-boundary match. Test names containing 'honour' (identifiers) are left as is.

Tested: typos, ruff.

* docs: trim the getting-started MLX route

The MLX section duplicated the backends guide's end-to-end and precision sections; it is now install, AMICA(backend="mlx") fit/transform, an AMICAICA one-liner and the float32 note, linking to the guide for the full workflow. The install heading no longer uses a bare GPU abbreviation.

Tested: mkdocs build --strict; linked anchors checked.

* docs: harness status and test working directory

progress_summary.md's harness line now matches AGENTS.md (every backend, #315), and the testing guide explains the conftest session chdir and how a test resolves repository paths.

Tested: mkdocs build --strict.

* docs: one description of reference resolution

default_reference_binary's docstring is now the single full description of the reference-binary resolution order; the --native-engine help summarizes it and points there. The native-backend page runs the harness with uv run.

Tested: --help renders; harness tests (32 passed).

* test: two-iteration fits for the map gate

The NumPy-vs-torch sensor-map gate (1e-9) failed on the macOS arm64 runner at 1.3e-9: the backends start within 1e-13 and EM amplifies the gap about sixfold per iteration, so five iterations was platform-sensitive. Two iterations keep the same gate with a 1e-12 gap while the pre-fix transposed maps still differ by 2.6-3%. Tested locally: module passes.

---------

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Write and read the EEGLAB sphere column-major (#336)

* Fix S-order tests to expect column-major (#336)

* Add cross-backend sphere export round-trip test (#336)

* Document the sphere-order fix (#336)

* Add native-binary sphere orientation test

* Compare MLX sphere bytes bit-exactly

* Fix sphere-order comments in write_amicaout

* Warn about old asymmetric-sphere output

* Correct scratch-history note on S order

* Polish sphere-order wording in docs/tests
* Add seeded native-binary oracle helper

Writes a pamica state in the reference's load_* formats, runs the
pinned v0.3.3 binary and reads its raw output arrays back, for
element-wise oracle tests (issue #333). Used by the tests that follow.

* Normalize component rows in doscaling

Each model's stored A block is the transpose of the reference's
mixing matrix, so a component is a block row. doscaling normalized
stored columns, which is not a change of scale; torch, NumPy and MLX
now rescale each component row with the matching mu/beta, exactly as
amica15.f90:1843-1851 does (issue #333).

Tested: invariance, cross-backend and doscaling=False byte-identity
tests; seeded native oracle matches A/mu/sbeta to round-off. Tests
that encoded the column rule are updated.

* Retune overshoot recipes for the new trajectories

With components rescaled, the aggressive-Newton recipes run monotone,
so keep_best restores never fired (failures or silent skips). Add
newtrate=3.0 and longer budgets where needed, move the pdftype=1
switch to kurt_start=6, pick n_mix=1 for the default min_dll stop,
and check the keep_best recompute only after a restore.

Tested: every retuned test passes and no longer skips.

* Retune two MLX cross-backend test configs

The variance-order guard rejected the 30-iteration fit (gap 5e-4), so
run 50 (gap 2.4e-3). One two-model sample sat on the rejsig=2.0
threshold between float32 and float64; use rejsig=2.5.

Tested: both tests pass for one and two models.

* Count scalestep from 1 and validate it

The reference parses scalestep but rescales every iteration. Keep it
as a pamica extension counted from 1 (iterations s, 2s, ...), so the
default 1 is the reference, and reject a non-integer or < 1 value at
construction on every backend instead of dividing by zero mid-fit.
Record it as row 13 of the differences page.

Tested: cadence and validation tests on torch, NumPy and MLX.

* Document component orientation in ADR 0006

ADR 0006 states the storage convention (components are rows of each
stored block), the doscaling fix and the Phase 8 layout change, and
amends ADR 0001; the issue #24 notes point to it. The changelog
records the behavior change with the oracle numbers.

* Make keep_best-under-reject tests non-vacuous

Both the MLX and torch tests ran monotone fits, so final_ll_ equal to
the last LL proved nothing. Use the overshoot recipe, require that
it restores without do_reject and that the do_reject trajectory also
ends below an earlier peak, then assert no restore.

Tested: both tests pass; the old configs were monotone.

* Check reject agreement up to float32 rounding

Exact agreement of the rejected set is config luck: float32 and
float64 per-sample LLs differ by up to ~0.1 by the second pass.
Record each pass, require every disagreement to lie within the band
measured on that pass, cap the band and the disagreement count, and
add controls where a threshold or trajectory divergence fails.

Tested: seeds 7 and 8, one and two models, at rejsig 2.0.

* Note ADR 0003 figures predate the row rule

* Rebuild paper.pdf

* Derive MLX pin tolerances from float32 study

* Check variance order against a float32 band

* Band-check the pdftype1 reject test

* Use semantic line breaks in ADR 0006

* Note Phase 8 aligns the initial A norm

* Word pre-fix norm ranges as history

* Note parity rows predate the row rule

* Document when scalestep is validated

* Say oracle maxima span backends and models

* Mirror the zero-norm guard in Newton test

* Test the rescale on Newton-reached states

* Rescale a permuted comp_list across backends

* Test zero- and NaN-norm rows stay untouched

* Fail, not skip, on recipe preconditions

* Never read stale or partial oracle output

* Check out full history in pytest CI jobs

* Fail byte-identity pins in CI, not skip

* Use the CI-proven overshoot recipe in torch

---------

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Count schedule gates from 1 like the reference

Route every iteration-schedule gate in the torch, NumPy and MLX backends through a shared pamica/schedule.py. The Newton switch (iter .ge. newt_start), the numdecs reset on the switch-on iteration, the iter > newt_start ceiling ratchets, the outlier-rejection schedule and NumPy's restart-on-NaN window compared the 0-based loop index with the reference's 1-based setting and fired one iteration late. Share, kurtosis, A-freeze and writestep/histstep gates were already 1-based and now use the same helpers. Seeded against the v0.3.3 binary (newt_start=3, doscaling off, 6 iterations) the LL deviation drops from 1.2e-3 to 5.4e-11 (torch) and 2.1e-12 (NumPy).

* Update schedule-dependent tests to 1-based gates

Tests whose configurations or replays encoded the old 0-based newt_start/rejstart: the aggressive-Newton keep_best recipe moves from newt_start=1 to 2 (the exact pre-change trajectory its documented measurements came from), the LLt reject tests fire on iteration 6 of a 6-iteration fit, and the MLX ratchet/reset replays spell out the reference's 1-based arithmetic. Tested: the 11 affected files pass (179 passed, 1 skipped).

* Test schedule gates against 1-based iterations

Cross-backend real-fit tests (torch, NumPy, MLX): the first Newton step is iteration newt_start, the rho-rate ceiling ratchet opens only after it, the switch-on counter reset follows the reference's replay, rejection first fires on iteration rejstart, and NumPy's restart-on-NaN window covers the first restartiter iterations. Tested: 15 passed; against the pre-fix backends 10 of them fail.

* Add gated native oracle for the Newton start

Seeds the pinned v0.3.3 binary with pamica's own init (indir/load_*), newt_start=3, doscaling off, single-threaded, 6 iterations. Torch and NumPy match its LL trajectory and A to round-off (torch LL 5.4e-11, A 4.2e-10; NumPy LL 2.1e-12, A 1.1e-10) where the pre-fix code deviated by 1.25e-3 and 3.1e-2. A no-Newton control shows the compared window is Newton-sensitive. Tested with AMICA_RUN_FORTRAN=1: 3 passed; 2 fail on the pre-fix backends.

* Document the 1-based schedule gates

Changelog entry for issue #335: which gates moved, the native-oracle numbers before and after, the default-path cases that change, and how to reproduce an old trajectory. No row in docs/guides/amica-differences.md changes: the anchored share A-freeze is already recorded there.

* Rebuild paper.pdf

* Validate iteration-schedule settings

One shared validator, schedule.validate_iteration_setting, called by all three constructors with identical messages: newt_start must be an integer >= 0 whether or not do_newton is on (it also gates the rho-rate ratchet), and rejstart an integer >= 1 when do_reject is on (with 1-based counting, rejstart <= 0 skipped the reference's unconditional first pass; NumPy accepted 0, torch and MLX any value). NumPy also validates restartiter >= 0, maxrestarts >= 0 and histstep >= 1 under do_history (histstep=0 was a bare ZeroDivisionError mid-fit), and documents that restartiter=0 disables restart-on-NaN. schedule.every and periodic_due raise ValueError for a non-positive interval. Docstrings: newton_active cites amica15.f90:1666, newt_start=0 and 1 are stated to fit identically, and the MLX newt_start/newtrate entries document the newtrate ratchet. Tested: the new validation, restart-window and interval tests pass (36 passed with the reject suites).

* Reflow MLX schedule docstrings

The share_start/share_iter and kurt_start entries no longer break mid-clause. No code change.

* Record NumPy restart-on-NaN difference

Row 14 of the differences guide: the reference allows maxrestarts+1 restarts and, because startover is never cleared (amica15.f90:1046), never calls update_params again after the first one (:1115-1122); the NumPy backend allows maxrestarts restarts and resumes fitting. PyTorch and MLX have no restart-on-NaN path and stop with nan_ll.

* Correct the schedule changelog entry

Use the measured pre-fix deviations (1.25e-3 over 6 iterations, 1.33e-3 over 100) in the changelog and the oracle test docstring, list restartiter among the settings that now count from 1, and add the new validation.

* Test schedule gates at their boundaries

The arithmetic test gains the =1 rows for every helper (newton_switches_on at 1 and 0, past_newton_start at newt_start and newt_start+1, rejection with rejstart=1/rejint=1 through maxrej exhaustion, every/periodic_due at 1, the restart window at 1 and 0). The first-Newton-step test runs at newt_start 1 and 5, and a new real-fit test on every backend shows newt_start=0 and 1 give bit-identical trajectories on an overshooting run whose ratchets fire. Tested: 12 passed.

* Make the switch-on reset test discriminating

The replay-based reset tests passed with the reset one iteration late too. The cross-backend test now uses a searched configuration (4096 frames, lrate=0.8, newt_ramp=1, maxdecs=3, newt_start=3) whose likelihood decreases on iterations 2, 3 and 4, where the reference rule predicts no ratchet and the one-late rule predicts one; the fitted lrate ceiling must match the reference. It fails on all three backends with only newton_switches_on reverted and on the pre-fix backends. The MLX twin, which could not discriminate, is removed in favor of it. Tested: 3 passed; 3 failed with the reset reverted.

* Test the rho-rate gate at newt_start=20

A searched natural-gradient configuration (4096 frames, seed 4, lrate 0.5, maxdecs 2) completes a maxdecs cycle on iteration 21 on all three backends, so the rho-rate ratchet gate is now pinned at the shipped default: open for newt_start=20, closed for 21. This is the one default-path change of issue #335. Tested: 3 passed; 3 fail on the pre-fix backends.

* Document the reset-test configuration search

States the full search that selected the switch-on reset configuration, including the two-model hit found after the first commit.

* Shift Phase 7 recipes to 1-based settings

Phase 7 retuned its overshoot recipes under 0-based gating. With newt_start and rejstart counted from 1, n becomes n+1 so each documented trajectory reproduces bit for bit: the min_dll overshoot recipe (newt_start 2, newtrate 3.0) again peaks at 70 and stops at 71 on torch (1.9e-4 below the peak) and at 63/64 on MLX (4.2e-4); its do_reject variants move rejstart 5 to 6; the doscaling-rows NEWTON state uses newt_start=2; the min_dll-reachable test uses 6 and the slow rholrate ratchet test 51. Tested: the affected suites pass.

* Route scalestep through the schedule module

Phase 7's inline (iteration + 1) % scalestep rule becomes schedule.every in all three backends, and its scalestep check becomes schedule.validate_iteration_setting (same semantics and message; still only when doscaling is on), which drops the numbers imports it needed. Tested: test_doscaling_rows passes.

* Re-search the newt_start=20 gate configuration

Phase 7's component-row doscaling moved the old configuration's second ratchet off iteration 21, and the test's precondition guard failed as intended. The re-searched configuration (4096 frames, seed 1, lrate 0.6, maxdecs 3, newt_ramp 1) decreases at ll indices 4, 5, 8 and 9, 19, 20 on all three backends (smallest decrease 4.7e-4), ratcheting at 8 and 20. Tested: 3 passed.

* Note the combined doscaling and Newton-start parity

Changelog: with doscaling on (component rows, issue #333) the seeded 100-iteration Newton comparison stays within 3.9e-6 (torch) and 7.0e-6 (NumPy), where the #333 fix alone left 2.1e-4; scalestep now uses the shared schedule helper.

---------

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Add a loader for pre-change pamica in tests

pamica/tests/pre_change.py imports the whole package at a pinned commit
from git under a private name, so byte-identity tests can compare the live
backends with the pre-change ones in one process. It fetches only when the
object is missing (a --depth 1 fetch in a full clone makes it shallow), and
fails under CI but skips locally when the commit is unreachable.

* Store mixing matrix components as rows

All three backends now hold A as (n_comps, n_channels): row k is the
reference's column A(:, k), and model h's block is A[comp_list[:, h], :],
the same matrix as before, so W = inv(block) is unchanged (issue #334).
comp_list now indexes components in A as it does in the densities, so:

- the share metric compares pinv(sphere) @ A.T, the components' scalp
  maps, not stored columns of each block;
- a merge ties two components' mixing vectors and densities, and the
  gm-weighted dAk/zeta step averages one component across models;
- doscaling rescales each component row once, shared or not;
- accessors, W, the EEGLAB export (reference layout), load_results and
  the viz helpers follow the new layout.

Fits without a merge are byte-identical on every backend; ndtmpsum moves
by float round-off (it sums per component row, like the reference).
Existing sharing, doscaling and MLX tests are updated to the row layout;
two short sharing recipes are retuned because early scans now merge
most components. Tested: full non-slow suite and slow tests.

* Version saved models for the row layout

The PyTorch state_dict is now format_version 4 and the MLX save format 2.
A version 3 (PyTorch) or 1 (MLX) payload is converted through the shared
pamica.component_layout.rows_from_legacy_columns: without loss when its
comp_list is unmerged, and refused with a request to refit when a
component id is shared across models, since that merge was made under
the old column semantics. AMICA.save files (still version 2) go through
the same path.

Tested with test_component_rows_persistence.py (24 passed): round trips
in the new formats, and payloads written by the pre-change classes
loaded from git, converted or refused.

* Refuse a stale multi-model A in load_results

load_results now checks that each model's block of A inverts the W
written beside it (max |A_h W_h - I| <= 1e-3; measured at most 3.8e-7
for float32 MLX exports and 1.1e-15 for float64). A multi-model
directory written before issue #334 stored A with components as columns
and reads back scrambled (residual 1.5); it is refused with the remedy
instead of handing the viz helpers wrong sensor maps. Single-model
files are the same bytes in both layouts and still load.

* Test component rows and sharing semantics

test_component_rows.py pins, on the sample EEG:
- byte identity of every unshared configuration against the pre-change
  package (3627a6e) on PyTorch, NumPy and MLX: 1-3 models, doscaling,
  Newton, pdftype 0/1 and do_reject;
- the metric compares get_sensor_mixing_matrix columns, a planted
  duplicate merges, and a merged group shares one mixing vector;
- the EEGLAB A file is the reference layout, single-model exports are
  byte-identical, and load_results refuses a pre-change multi-model A;
- opt-in (AMICA_RUN_FORTRAN=1): updates from a merged load_comp_list
  state match amica15mac to float64 round-off on PyTorch and NumPy,
  with doscaling on and off.
All pass locally, including the gated oracle.

* Document the component-row layout

ADR 0007 records the row layout, the sharing semantics that follow from
it, the persistence and export policy, and the measured evidence; ADR
0006 and 0001 are amended, and ADR 0006's initial-A normalization line
now points to issue #341. The changelog, the differences page (sharing
rows, a new section for issue #334, and row 15: the reference's
single-precision density normalizers), and the backends, EEGLAB,
validation and API pages describe the new behavior. mkdocs build
--strict passes.

* Reflow two sharing test docstrings

Docstring text only; no test logic changes.

* Skip MLX wrapper routes without MLX

The merged-save refusal test's wrapper route fitted an MLX model without first checking MLX was installed, so it failed on the Linux runner. _wrapper_fit now performs the check for every wrapper route. Tested with MLX hidden: the new test files pass or skip.

* Pin byte identity to the Phase 9 epic head

The pre-change package is now 0930c0e, the epic head with Phase 9's schedule gates and without this phase, so the two sides differ by the layout change alone. The full byte-identity matrix (Newton included) and the persistence tests pass against it: 92 passed, 2 gated oracle tests skipped.

* Stop tests from fetching pinned commits

Tests that load historical code ran `git fetch origin <sha> --depth 1`,
which in a full clone records a shallow boundary that gc can prune
behind (issue #343). pre_change.py now only reads: it checks the pin
with `git cat-file -e` and, when the commit is missing, fails under CI
and otherwise skips with the command that fetches it (`git fetch origin
<sha>`, or `git fetch --unshallow origin` in a shallow clone). Pins must
be full SHAs.

test_mlx_fit_noop.py, the last loader that fetched, now imports the
whole package at its pin through the same helper; its fits are still
bit-identical. test_pre_change_loader.py runs the loader behind a git
wrapper that logs its arguments and execs the real git, and asserts no
fetch ran and the shallow state is unchanged (a negative control with a
real fetch trips the assertion). Tested: loader and fit-noop tests pass,
also with CI set.

* Record the single-precision density constants

The reference writes its density normalizers as dble(<literal>), which
gfortran reads in single precision: log(dble(1.772453851)) in the
rho == 2 branch is 3.0e-8 above log(sqrt(pi)) (measured in the seeded
oracle), and the Gaussian and cosh families' normalizers differ from
pamica's double values by 3.7e-10, 2.0e-8 and -2.1e-8 (computed).
Row 16 of the differences page records this, citing issue #344, which
decides whether to adopt the reference's values. The torch and MLX
comments that called _LOG_SQRT_2PI and _LOG_NORM_COSH_* bit-for-bit
with the binary are corrected, and the merged-state oracle's
maxrho=1.99 cites the issue. Comments and docs only.

* Show early mass merges happen in the reference

With share_comps on, the pinned binary's scan runs but never merges:
Spinv2 is never allocated, so every similarity is NaN, even at
comp_thresh=0. Its similarity (Spinv2 = Spinv^T Spinv) applied to its
own state after 8 iterations from pamica's initialization merges exactly
the pairs the torch and NumPy scans merge (32, 32 and 30 at 0.9, 0.95
and 0.99). In the recipe that went non-finite, the reference's update
from pamica's post-scan states collapses in step: the second model's
gm falls to 5.04e-3 on both after the first scan, and after the second
reaches zero in one iteration and NaN in the next.

Two gated tests pin this (AMICA_RUN_FORTRAN=1, both pass). The
differences page, ADR 0007, the changelog, the validation table and the
torch/MLX comments that called the scan unrunnable now say what it does.

* Harden the pre-change loader and its guard

pre_change.py now fails under CI, and otherwise skips, with its own
message when git is not installed (instead of a raw FileNotFoundError)
and when the tests do not run from a git checkout (instead of advising
a fetch that cannot help). The repository root is a parameter, and a
cached alias must map to the same commit or loading raises.

The test's PATH wrapper now refuses fetch, pull and every depth or
shallow option without running git, so the committed negative control
(a fetch sent through _git) is caught by _assert_no_fetch and leaves the
clone untouched. New tests cover missing git, a non-checkout root and
alias reuse. Tested: 10 passed, also with CI set; the clone is not
shallow afterward.

* Validate comp_list dtype in the layout conversion

rows_from_legacy_columns now refuses a non-integer comp_list (float or
NaN ids) with its own ValueError before the bounds check, instead of
letting it reach the indexing. test_component_layout.py adds direct
unit tests on small arrays for every ValueError branch: the comp_list
rank, the dtype, the A shape, an out-of-range id, an id repeated within
a model, and an id shared across models (the refit message), plus a
block-for-block conversion of an unmerged, permuted comp_list.
Tested: 11 new tests pass; persistence tests unchanged.

* Refuse a truncated A in load_results

load_results now checks that the A file holds num_comps * nw values
before reshaping it to (num_comps, nw), and names the file and both
counts when it does not. Before, a length no reshape accepts surfaced
as a bare NumPy reshape error, and one short by a whole component
reshaped to the wrong width and failed later as a matrix-product error.
Tested on truncated copies of a real two-model export (both cases).

* Raise instead of assert in NumPy _write_results

The fitted-A and outdir preconditions of AMICA_NumPy._write_results are
now explicit RuntimeErrors, since assert is stripped under python -O.
The comment above them also said the reference omits A from its
output; the binary does write it, in the layout pamica now writes.
Tested: NumPy write and load tests pass.

* Widen the byte-identity matrix

The unshared byte-identity test now also covers the Gaussian, logistic
and sub-Gaussian cosh families (pdftype 2, 3 and 4, single model, n_mix
1 for family 4), four blocks of 1024 samples, a keep_best fit whose
log-likelihood falls so the safeguard restores an earlier snapshot
(asserted), and best-of-two restarts, on every backend that supports
each. The NumPy backend has only the generalized Gaussian and no
keep_best, so those cases run on PyTorch and MLX; the test says so and
test_numpy_rejects_other_pdftypes pins the first reason. The helper now
lets a configuration override its defaults, and compares final_ll_.
Tested with CI=1: 66 passed.

* Check the MLX double conversion on save again

test_old_mlx_save_converts_losslessly now saves the converted model
again, checks the file is format 2 with a component-row A, reloads it,
and compares A and every output with the saved model, as the torch test
does: a second load must not convert the already-converted A again.
Tested: both recipes pass.

* Fix the Spinv citation and sharing comments

The torch _identify_shared_comps docstring cited Spinv's allocation at :551; it is amica15.f90:569. A test_mlx_newton docstring said merged-away columns and curvature not indexed by mixing column; with components as rows the words are component and mixing vector. Comments only.

* Label the merged-oracle figures consistently

ADR 0007 and the changelog now say the 3-iteration oracle figures are the worst of doscaling on and off, and give the old column semantics' error as the matching worst value (A 0.21, LL 4.3e-4, not 0.20 and 4e-4); the test comment lists both measurements. Docs and comments only.

* Describe component-row sharing in the notes

AGENTS.md and .context/progress_summary.md described the old column metric and merge as working. They now describe the component-row semantics (ADR 0007, #334): the metric compares de-sphered component mixing vectors, the merge ties components, and the reference binary's own scan never merges because Spinv2 is never allocated. Docs only.
* Reject unknown AMICA_NumPy keyword arguments

AMICA_NumPy(**kwargs) forwarded every keyword into a params dict read
with params.get(...), so a typo (max_iters=50) or an option the legacy
backend does not implement (keep_best=True) constructed silently and
had no effect.

__init__ now validates **kwargs before any other processing: an
unrecognized name raises TypeError with a difflib-based suggestion,
and a name implemented on the PyTorch backend (AMICATorchNG) but not
this one raises TypeError naming it and pointing to
AMICA(backend='torch'). The accepted-keyword set is derived from
_CONSUMED_KEYS/_CANONICAL_TO_NUMPY_KEY (the same source the
params-file routing uses) plus this constructor's own named
parameters, so the two surfaces cannot drift apart. The canonical
spelling of the three renamed settings (min_nd/maxdecs/share_iter) is
now also accepted as a keyword argument directly, translated the same
way a params file's setting already is; passing both spellings at
once raises TypeError.

* Add tests for AMICA_NumPy kwargs validation

Real construction only, no mocks (no .fit() needed to exercise
__init__'s validation): a typo raises TypeError with a suggestion; an
unsupported cross-backend option raises the specific message; the
three canonical-spelling aliases are accepted and their conflict is
rejected; every documented NumPy option constructs, alone and all
together; and a source scan proves _CONSUMED_KEYS matches every
params.get(...)/raw_params.get(...) read site exactly, so the
accepted-kwargs set cannot silently drift from what __init__ actually
reads.

* Document AMICA_NumPy kwargs rejection behavior change

docs/changelog.md: Unreleased entry noting the legacy NumPy backend
behavior change (unknown or unsupported options now raise instead of
being ignored). docs/api/numpy-backend.md: a short paragraph mirroring
the existing pdftype callout.

* Improve AMICA_NumPy kwarg validation (review)

PR #347 review follow-up, several items together since they touch the
same functions:

- files/data_dim/field_dim now work as constructor keywords, with the
  same meaning as in params_file, instead of passing the accepted-
  kwargs check but staying silently inert (fit() with no data still
  said "no 'files' configured" even after a caller set one via a
  kwarg). Conflicting values between params_file and a kwarg raise;
  the same value from both does not.
- n_models/n_mix (AMICATorchNG's spelling of this backend's own
  num_models/num_mix) get a message naming the correct spelling
  instead of the generic (and here wrong) torch-only message.
- n_channels gets its own message: it is inferred from the data
  passed to fit() on every backend, not a constructor keyword on any
  of them (the AMICA wrapper's own fit(**kwargs) excludes it too).
- Every offending keyword in one call is named in a single error,
  however many different categories are mixed together. A first
  version raised on the first category found, so
  AMICA_NumPy(n_models=2, totally_bogus=1) named only n_models and
  dropped totally_bogus.
- _torch_only_options() imports AMICATorchNG lazily, inside its own
  body, called only on the error path, so numpy_impl carries no
  module-level dependency on torch_impl.
- Groups difflib/inspect with the other stdlib imports.
- Hardens the source-scan test: fails loudly on a params[...] bracket
  read or a non-literal params.get(...) key, either of which would be
  invisible to the scan.
- Adds typo-suggestion tests for params_file/use_tqdm alongside verbose.

* Fix dropped keyword in AMICA.fit mlx check

PR #347 review item 4: the same single-category drop
AMICA_NumPy._reject_unknown_kwargs had (issue #346) existed here too --
backend='mlx' with fit(dtype=..., blocksize=...) named only 'dtype'
(ValueError) and silently dropped 'blocksize' from the message, since
the torch-only check raised before the unknown-keyword check ever ran.

Both categories are computed before either is raised; a combination
now names both in one TypeError. The two single-category messages are
unchanged (still ValueError for torch-only alone, TypeError for
unknown alone), so existing callers/tests keep working.

* Document NumPy kwargs review follow-up

docs/changelog.md: a Review follow-up sub-bullet under the Phase 14
entry. docs/api/numpy-backend.md: mentions files/data_dim/field_dim as
keyword arguments and the n_models/n_mix/n_channels-specific
messages.
* Split the M-step into direction and update

Compute the natural-gradient/Newton step and its norm (the reference's
dAk and ndtmpsum) in _update_direction and apply it in
_update_parameters, in all three backends, sharing one set of
intermediates. The fit loops call both back to back, so every fit is
byte-identical (checked on 8 configurations per backend).

* Follow the reference's iteration order

Run the likelihood-decrease response and the stopping checks before the
parameter update, and exit on a stop before any parameter moves, as
amica15.f90:1015-1122 does, in all three backends. The step taken right
after a decrease now uses the halved rate, and a stopped fit returns the
parameters whose likelihood is ll_history[-1]. Keep the reference's
working rho rate and its ceiling apart (rholrate, rholrate_cap), reset
inside the A-update branch. Fits without decreases or stops are
byte-identical; seeded oracle deviations drop from 1.4e-2 to 2.7e-5.

* Hold A on the reference's schedule without sharing

Apply the reference's A-update guard (iter >= share_start and
mod(iter, share_iter) <= 5, amica15.f90:1803) in every fit, share_comps
on or off, through schedule.share_freeze, replacing the anchored
post-merge window. Every backend now rejects share_iter < 7 (NumPy:
share_int) whether or not sharing is on, since a shorter cycle would
never update A again. Default fits of 100 or more iterations change.

* Add iteration-order tests and native oracle

Always-on cross-backend tests for the freeze iterations, the decrease
timing, stop semantics, keep_best pairing and share_iter validation, plus
an opt-in comparison against the native binary (AMICA_RUN_FORTRAN=1).
All pass; the native comparison fails on the pre-change code.

* Retune tests to the reference's iteration order

Existing tests that pinned the old order (update before the checks, a
freeze gated on share_comps and anchored on share_start, one rho rate) now
pin the reference's, or use recipes that still exercise what they test.
Full non-slow suite and the gated native oracles pass.

* Document the reference's iteration order

ADR 0008, changelog entry with oracle and parity numbers, and the
differences guide: the A-freeze is now the reference's arithmetic, and
convergence stops return the parameters final_ll_ describes.
mkdocs build --strict passes.

* Test the rho-rate ceiling as its own value

Rate tests that read rholrate as the ceiling now read rholrate_cap, the
cross-backend state copies carry it, and a new test round-trips it through
state_dict on PyTorch and MLX, including a save without the key.
Affected tests pass, including the slow torch ratchet test.

* Stop overshoot recipes on the first decrease

The keep_best overshoot recipes stopped via min_dll=1e-4/maxincs=2,
which under the reference order ended at the peak or below it depending on
round-off: at the peak on the macOS CI runner, and in 1 to 2 of 12 data
perturbations locally. maxincs=0 with min_dll=1e-8 stops them on their
first likelihood decrease instead, which held in all 12 perturbations on
PyTorch and MLX, with and without rejection. Affected tests pass.

* Stop identically on non-finite values

Every backend now keeps a non-finite likelihood out of its history (NumPy
recorded it), stops before applying a non-finite step or gradient norm
(new degenerate stop_reason nan_direction), and stops on non-finite
parameters right after an update (nan_params, now also in PyTorch and
NumPy). NumPy defers only an A/W corruption inside its restart window to
restart-on-NaN, which can repair it. Cross-backend injection tests pass.

* Clear numincs on a NumPy restart

The reference's min_dll comparison with the NaN likelihood of its restart
iteration is false, so it zeroes numincs (numdecs is left as it was).
NumPy's restart carried the small-gain count across; it now clears it.
The new test stops at index 8 as expected and at 6 without the fix.

* Validate saved learning rates on load

from_state_dict (and the MLX file load) now refuses a missing or
non-finite lrate, lrate_cap, newtrate, rholrate or rholrate_cap with a
ValueError naming the field; a missing rholrate no longer surfaces as a
bare KeyError on PyTorch. Tested on PyTorch and MLX.

* Validate share_start on every fit

The A-freeze reads share_start whether or not share_comps is on, so every
backend now requires an integer share_start >= 1 always. The share_iter
message now says why a value below 7 is refused (A held permanently from
share_start), and NumPy's names both of its spellings. Both checks are
reached through a Fortran params file on all backends; tests pass.

* Test every convergence stop's parameters

The stop test now runs min_dll, grad_norm, grad_norm_floor and lrate_floor
on every backend, each with a precondition on its stop_reason and a twin
that runs the same iterations to max_iter. All 12 cases pass.

* Test that a NumPy restart skips the update

A pass-through recorder on _update_parameters shows the restarting
iteration applies no update; the same fit of the pre-change code, loaded
through pre_change.py as a control, updated on it. Passes.

* Test MNE n_iter_ after each kind of stop

to_mne_ica().n_iter_ equals iteration + 1 and len(ll_history) after a
min_dll stop and after a max_iter fit, on PyTorch and MLX. Passes.

* Take the oracle floor over two thread pairs

The reference noise floor is now the larger deviation of a 2- or 4-thread
binary run from the 1-thread run. The multi-threaded runs are not
reproducible, so the docstring gives each floor's range over three runs.
All four gated oracle cases pass in each run.

* Fix docstrings and Fortran citations

mir_history_ records no waypoint on a stopping iteration; MLX documents
its rholrate/rholrate_cap pair as PyTorch does; NumPy fit() gets the
iteration-order notes. Rejection is cited as amica15.f90:1136-1140 and the
check ranges start at 1056 and 1019. Docstring and comment changes only.

* Document the non-finite stops and validation

ADR 0008 states the final_ll_ invariant with its keep_best qualifier and
the review's changes; the changelog and differences guide cover the
non-finite stops, share_start validation and saved-rate checks, with the
two-pair oracle floors. AGENTS.md and the context notes describe the
unconditional A-freeze. mkdocs build --strict passes.
* Use the reference's single-precision constants

The reference writes its density normalizers as default-kind literals
widened with dble, so the binary uses their float32 roundings. All three
backends now read those values, and epsdble's, from one shared module,
pamica/reference_constants.py; MLX holds the rho == 2 normalizer in its
lgamma table. The literal-formula tests and the benchmark's copy now use
the reference's values. Tested: pdf-family and backend test files pass.

* Compare pre-change fits at the old constants

The byte-identity tests against code from before issue #344 (component
rows, doscaling off, the MLX pre-Phase-3 tip) now give the live backend
its old density constants through pre_change.use_pre_344_constants, so
they still isolate the change each was written for: their two-model and
non-GG fits reach the changed normalizers. Tested: the three files pass.

* Pin the constants against the reference

test_reference_constants.py checks each shared constant against the
float32 rounding of its literal on the cited reference line (exact
rational arithmetic), pins the sweep of amica15.f90 and its header,
checks that no backend keeps a copy, checks every backend's rho == 2
log-density, and (opt-in) seeds the native binary for pdftype 2, 4 and
1. Tested: 23 passed with AMICA_RUN_FORTRAN=1.

* Run the merged oracle at rho == 2

The merged-state oracle no longer holds maxrho at 1.99: its warm seed
now has mixtures clamped at 2, it asserts that, and the code before
issue #344 is replayed from the same state as a control (off by 2.8e-9
in log-likelihood after 1 iteration, against 2.2e-15 now). Tolerances
unchanged. Tested with AMICA_RUN_FORTRAN=1.

* Document the single-precision constants

Changelog entry with the oracle numbers; the differences guide gains a
section on the constants, and row 16 now records the one kind of
single-precision literal pamica keeps at its decimal value (defaults of
input.param keys); the validation guide notes the new family oracle.
Tested: mkdocs build --strict, typos.

* Note the reference's unset lrate0 in row 16

Without an lrate key the reference never sets its ceiling lrate0, so
its lrate drops to 0 after the first iteration (checked against the
pinned binary). Row 16 now says so and uses comp_thresh as the example.
Tested: mkdocs build --strict, typos.

* Skip MLX before patching its constants

On a host without MLX the byte-identity tests imported the MLX module
to patch it before the check that skips them. Resolve the classes
first. Tested: the two files pass with MLX and skip its cases without.

* Remove the unused compute_log_pdf

numpy_impl.pdf.compute_log_pdf had no callers in the package, tests or
docs. Tested: the pdf and viz tests pass.

* Plot the fit's density in compute_pdf

compute_pdf, which viz.plot_pdf_fits draws, now takes its normalizers
from pamica.reference_constants, including the reference's
single-precision sqrt(pi) at rho == 2, so it draws the density the fit
uses. Its docstring says it is a plotting helper with the legacy
pdftype numbering. Tested: the pdf, viz and constants tests pass.

* Scan every backend module for copies

The no-copy test now scans every .py file of torch_impl, numpy_impl and
mlx_impl with comments and strings blanked, and flags the linear forms
(sqrt(pi), sqrt(2*pi), ...) as well as the log forms and the literals.
A new test checks that it flags every definition this change removed.
Tested: passes; fails on the old pdf.py sqrt(pi) line (checked by hand).

* Assert the family held in the oracle

The family oracle now asserts that every source kept its family and no
kurtosis switch ran (n_kurt_done == 0), for pdftype=1 with num_kurt=0
too, in the live and the pre-change runs. Tested with
AMICA_RUN_FORTRAN=1.

* Check the pre-#344 constant substitution

A control test: every constant use_pre_344_constants patches differs
from the live value, the patch takes effect, it covers every changed
constant each backend module reads, and on MLX it restores the
pre-change rho == 2 table entry. Tested: passes on all three backends.

* Note the plotting helper in the changelog

compute_pdf now draws the fit's density and compute_log_pdf is gone.
Tested: mkdocs build --strict, typos.
* Run the native binary from its own draw

run_drawn_reference runs the binary with nothing loaded and a fixed
seed; with max_iter=0 it writes its drawn initialization. The run and
read-back are shared with run_seeded_reference.

* Normalize the drawn initial mixing matrix

Every backend now draws its initial A through the shared
pamica.initialization.initial_mixing: per model, 0.01*(0.5-u) with the
diagonal set to one and each component divided by its norm, as the
reference does (amica15.f90:805-823), including the NumPy restart
redraw. A supplied or loaded A is used as is. Tested in later commits.

* Start pre-change classes from the new init

The byte-identity pins (fb13d76, 0930c0e, 2e04006) and the #339 oracle
control now fit the historical class from the live normalized initial A
through pre_change.with_normalized_initial_mixing, so each still isolates
its own change. Tested: all those byte-identity tests pass.

* Test the normalized initial mixing matrix

Always-on cross-backend start checks with a pre-change control, the
supplied/loaded A paths, the NumPy restart redraw, and two gated native
oracles. Tested: 34 passed with AMICA_RUN_FORTRAN=1.

* Re-search the MLX pdftype=1 restore recipe

The normalized initial A made seed 3 at lrate 0.5 climb monotonically
to max_iter, so the restore was never exercised. Seed 6 at lrate 0.4
(kurt_start 7) overshoots by 0.17 after all five switch passes, on all
12 perturbed inputs. Tested: test_mlx_keepbest.py passes.

* Re-record the MLX fit-path canary

The normalized initial A moved the recorded trajectory (ll_history[0]
by 1.55e-5). Pins re-recorded on the M4 Pro and tolerances re-derived
from the same 48-draw float32 perturbation study, floored at one float32
unit: no draw moved ll_history[0]. Tested: the canary passes.

* Seed the doscaling oracle from a conditioned state

From seed 42's normalized init, one sample sits 2.3e-7 from a mixture
center; the mu update amplifies the 1e-14 sphere round-off to 1.3e-9
(PyTorch and NumPy differ by 7.9e-10), over the 1e-9 bound. Seed 45
stays 7x under every bound; a new precondition asserts the backends'
own disagreement sits 3x below each bound. Tested with the binary.

* Re-search the early-merge collapse recipe

From its normalized initial A, seed 20's second model keeps gm 1.2e-3
after the second scan instead of collapsing to zero in one iteration.
Seed 23 (of 23, 32, 40, 41 in seeds 0-49) collapses as the test needs.
Tested with the binary (AMICA_RUN_FORTRAN=1).

* Document the normalized initial mixing matrix

Changelog entry with the measured behavior change, ADR 0006's remaining
difference marked resolved (and ADR 0007's pointer), the differences
guide's init wording and the re-searched collapse recipe. mkdocs build
--strict passes.

* Re-seed the MLX MIR restore test

From its normalized initial A, seed 0's discarded step moved the MIR by
2.1e-4 relative (3.4e-5 to 2.1e-3 over 8 perturbed inputs), inside the
1e-4 margin on the CI GPU. Seed 7 moves it by 2.3e-3 to 2.9e-3 on every
input. Tested: test_mlx_mir.py passes.

* List the re-seeded MLX MIR test in the changelog

* Refuse a zero-norm component in normalization

normalize_components now raises ValueError on a zero or non-finite
row norm instead of returning NaN. The reference path cannot reach it
(the unit diagonal keeps every norm >= 1). Tested: zero and NaN rows.

* Check a supplied NumPy A at fit start

A supplied A (or one left by a previous fit) now raises ValueError on
a wrong shape, a non-finite entry or a singular model block, through
the shared validate_supplied_mixing. Before, it raised LinAlgError deep
in the fit, or (NaN) ended with no stop reason. Tested: all three.

* Test NumPy fix_init and the wrapper round trip

fix_init (an attribute) starts from the stacked identity and leaves
the generator untouched; AMICA save/load restores the PyTorch A bit
for bit. Tested: test_initial_mixing.py passes.

* Separate the freeze from the collapse figures

The 9.0e-4 gm drop is pamica's own fit, whose A is held on those
iterations; the 4.2e-3 both sides reach comes from the freeze-free
continuation. Docs and test comment now say so. Docs only.

* Tighten three comments from the review

ADR 0006 marks the init difference as former; the _FLOOR_MARGIN
comment names the worse of doscaling on and off; pre_change notes a
component is a block row in both layouts. Comments only.

* Record the supplied-A check in the changelog

* Sphere both oracle sides with the reference S

The merged-state oracle now sphers pamica's data with the binary's own sphere, so both sides start from the same bits after the Phase 12 and 13 merge. Tested: gated component-row oracles pass.

* Start the pre-344 control from the new init

The pre-#344 control class predates #341, so it now starts from the normalized initial A the seeded reference uses. Tested: gated family oracles pass.
* Delete the dead legacy pdf code in numpy_impl/pdf.py

choose_pdf_type had no callers, and compute_pdf's pdftype 2-4 branches
were unreachable: every caller uses the generalized Gaussian. Tested
with the pdf, reference-constant and viz suites.

* Correct public docstrings against the epic code

Rewrap docstring lines that began with an issue number, which the API
pages rendered as headings; state the sphered-space shapes of the
wrapper's mixing and unmixing accessors; name restart_error among the
degenerate stops; drop the claim that natural gradient alone falls short
of the reference; fix the NumPy n_channels and doscaling wording.
Checked on the bundled sample (accessor shapes and composition).

* Say that maxdecs counts decreases, not consecutive ones

The reference resets its decrease count only at each ratchet and when
Newton switches on (amica15.f90:1062-1076), as every backend does.

* Describe the reference's iteration order in the concept pages

How AMICA works now walks one iteration as the code runs it: the E-step,
the likelihood-decrease response and stopping checks, the exit before
any update, then the update with Newton, the A-freeze windows and the
doscaling rescale. The false claim that every iteration raises the
likelihood is gone, and the flow figure shows the checks before the
update. What is AMICA covers the unit-norm scale convention.

* Fix the EEGLAB guide's scalp-map example and multi-model note

The variance-order example took scalp maps from the sphered-space
get_mixing_matrix; it now uses get_sensor_mixing_matrix. The multi-model
note still said the per-model layout differs from the reference's, which
#159 and #334 fixed. The guide index lists all four guides.

* Correct the backends guide on devices, CUDA float32 and scope

Auto device selection picks MPS first and the wrapper redirects it to
CPU for float64; the raw AMICATorchNG raises there instead (checked on
this host). CUDA float32 is about as fast as float64, not faster, as the
same page says above. The native backend is listed, the wrappers' two
backends and their parameter-file route are stated, and a stale 'sweep
is being finalized' note points to the published sweeps.

* Document the raw PyTorch backend's automatic device choice

With device=None it picks MPS first and, at the float64 default, raises
on an Apple machine; the AMICA wrapper redirects to CPU. Checked on this
host.

* Update install, accessor and precision notes on the entry pages

pamica is on PyPI (0.3.3), so getting started no longer says a release
is planned, and says unreleased changes need a source install. The
README's get_mixing_matrix comment called the sphered-space matrix
scalp maps; the examples now use get_sensor_mixing_matrix for maps. The
home page no longer calls float32 faster, the final_ll_ note covers the
max_iter case, and the NumPy CLI takes either parameter-file format.

* Correct the API overview and backend pages

List AMICANative in the public import surface; install extras with uv
from PyPI or a source checkout; correct the NumPy page's n_channels
claim (the raw PyTorch and MLX constructors take it) and list what that
backend lacks; drop the viz page's reference to the closed #159 as an
open question; replace em-dashes in list items.

* Say the reference compiles Newton off, not on

amica15_header.f90 declares do_newton = .false.; only the bundled
input.param turns it on. The Newton docstrings and the concept page
said the reference runs Newton by default.

* Correct the differences guide's Newton row and backend table

The reference compiles do_newton off (only the bundled input.param
turns it on); row 5 names singular_ll; the backend table gains the
variance_order, scoring and wrapper rows and the wrapper .pt save; an
obsolete pre-#334 comparison of two retired sharing metrics is replaced
by the current single-kernel account; the LLt exception names the last
iteration; table cells say none instead of an em-dash.

* Bring the validation guide's non-numeric claims up to date

The degenerate-fit row still said raw backends were ungated (#306 is
done) and named only the likelihood stops; the EEGLAB section still said
multi-model files are not in the reference layout. Parity figures are
left to #351.

* Make the Unreleased changelog one coherent account

Group the phase entries by what changed (fitting, MLX and the wrappers,
the MNE wrapper, persistence and exports, fixes, removals, docs) and
open with a warning that default fits differ from 0.3.3 and why. Drop
the entries a later phase superseded (Phase 8's pending #344 note, the
polish round's harness and Phase 1's remaining-gap notes), merge the
duplicate pdftype-default entries, correct the n_channels claim, and
mark measurements taken before later phases. No figure is changed.

* Map the shared modules and the epic's fitting changes in AGENTS.md

Add rank, schedule, initialization, component_layout and
reference_constants (with blocktune, restarts and the params reader) as
the shared-decision modules, plus mne_compat, native and metrics; name
amica15.f90 as the reference source; add a status line for the
reference-fidelity changes and the #306/#339 degenerate-fit extensions;
shorten the epic #278 history sentence.

* Bring the context plan and progress summary up to date

Record epic #324's phases and the reference-fidelity changes, the
#306/#339 degenerate-fit extensions and the #334 sharing oracle; close
the stale release items (repository transfer, docs deploy, PyPI and
Zenodo are done) and the #15 coverage item.

* Document the opt-in native-binary tests and MNE's pre-whitener

The testing page names AMICA_RUN_FORTRAN=1 and the mne_tests directory;
the MNE page's source formula includes the per-channel-type scaling the
export applies.

* Say shared component, not shared column, in AMICAICA's docstring

* Note that doscaling is on by default in What is AMICA

* Avoid em-dashes in the plan's updated phase headings

* Mention the post-update non-finite stop in How AMICA works

* Reword the paper's framing in nuanced terms

Drop negation-contrast and candor framing outside the Validation
section, count releases as of today (nine, six on PyPI), and describe
the MEG status as end-to-end use without measured parity.

* Drop candor framing in the differences guide and ADR 0007

* Rebuild paper.pdf

* State the complete transform in get_unmixing_matrix

The docstring called W @ get_sphere() the full transform; transform
also subtracts the mean before sphering and the model center after.
The MIR docstrings and the metrics page now call W @ sphere the linear
part. Checked on the bundled sample with two models: the full formula
matches transform exactly, W @ sphere @ X is off by up to 4.2, and MIR
is unchanged by a shift of the data.

* Say which default fits #335 and #334 reach, in the changelog

The opening warning said the schedule gates and component rows matter
only with Newton, rejection, restarts or sharing; they also reach plain
default fits in narrow cases (a maxdecs ratchet past newt_start, and
the per-component ndtmpsum sum near a grad-norm stop). The stopping
iteration's skipped steps are listed in code order (kurtosis switch,
share scan, MIR waypoint, rejection), in ADR 0008 too, and the #304
entry says the fit applies the file's settings.

* Replace the two em-dashes left in the validation guide

Prose only; no figure on those lines changed.

* Say that fit applies from_params_file's settings

from_params_file stores the file's settings on the instance; fit passes
them to the backend constructor as defaults and emits the warning.

* State effects directly in the prose this PR added

Reword contrast framings in the concept page, home page, viz page and
module docstring, and the differences guide's opening, per the style
rule: say what the code does and its known effect.

---------

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Take the wrapper's fit defaults from the backend

AMICA.fit restated max_iter/lrate/do_mean/do_sphere/do_newton, and its
lrate (0.05, runamica15.m's) differed from every backend's 0.1. It now
reads them from the selected backend's signatures, so AMICA and AMICAICA
fits without lrate run at 0.1. Tests for torch and MLX: resolved
defaults, a default wrapper fit equals the default backend fit, and one
set of constructor defaults across the backends.

* Move the MPS/float64 device fallback into AMICATorchNG

AMICATorchNG(n_channels) with default arguments raised ValueError on
Apple Silicon: device=None picked MPS, which has no float64. The
constructor now uses the CPU for that case and logs a warning; the
wrapper and AMICA.load pass device through and drop their copy. An
explicit device='mps' at float64 still raises. Torch-only (MLX has no
device choice). Tested on MPS: raw construction and fit, float32 auto,
explicit MPS, load with device=None.

* Remove the dead pamica/data package-data entry

No pamica/data directory exists. Built with uv build: the wheel's file
list is unchanged, it imports from a clean venv, and
numpy_impl/params.json is inside.

* Document pamica's defaults and the backend device choice

Add a defaults table to the differences guide: pamica against the
compiled amica15 binary and EEGLAB's runamica15.m, why pamica follows
the binary, and how to reproduce an EEGLAB run. Update the pages that
stated the wrapper's lrate or the raw backend's MPS raise. mkdocs build
--strict passes.

* Add the Phase 17 changelog entries

The wrapper lrate change joins the 'Default fits differ from 0.3.3'
warning; entries for the backend device fallback and the package-data
cleanup.

* Give AMICA_NumPy the shared default settings

params.json set do_newton on and max_iter 2000, the only shared settings
where NumPy differed from AMICATorchNG; it now sets them off and 100, as
does the constructor's max_iter fallback (reached by the CLI). Fits
shorter than newt_start are byte-identical. New test_default_settings.py
holds NumPy and MLX to every shared AMICATorchNG default (the MLX check
moves there from test_wrapper_backends.py).

* Correct the degenerate-fit test's measured stop

The comment said torch stops on nan_ll at iteration 3; both backends
stop on nan_params with iteration 2 in all 16 measured runs, at either
lrate. The test asserts only a degenerate stop.

* Document the NumPy backend's shared defaults

Defaults section, validation guide's stops table, NumPy API page, and a
changelog entry (NumPy joins the opening warning's defaults line).

* Give AMICANative pamica's shared defaults

The engine wrote the bundled input.param's values (lrate 0.05, Newton
on, max_iter 2000, block_size 512, ...). It now takes every setting the
binary has a keyword for from AMICATorchNG's signature through
FORTRAN_TO_PAMICA_KEY; binary-only knobs are unchanged. The default
block is pamica's 8192 samples split over max_threads and capped at the
data, since the binary's block_size is per thread and a block longer
than the data gives an all-NaN fit (#292). The native oracles now start
from the bundled input.param itself (byte-identical files). Tested with
the v0.3.3 binary: native, oracle and default-settings tests pass.

* Document AMICANative's shared defaults

AMICANative joins the defaults table's pamica column; the section and
the native API page cover the settings it cannot write, the per-thread
block size, and running the binary as an input.param configures it. The
changelog records the behavior change.

* Say where AMICANative's default pcakeep comes from
* Add issue-351 Newton independent-seed study script

* Add hallu driver for the Newton seed study

* Record harness, same-start and fixture parity runs

* Record share_comps counts and float32 agreement runs

* Update share_comps merge counts in differences guide

* Re-state the harness rows with same-start parity

* Add an MLX float32 option to the Newton seed study

* Mark the issue-27 and issue-51 records as pre-epic

* Update the harness figures in AGENTS and progress summary

* Add ensemble scripts and dimsweep LL records

* Add the hallu CUDA dimsweep LL record

* Re-state the cross-backend LL table with the chaos caveat

* Reword the harness and cross-backend notes

* Pin the ensemble lrate in reproduce_table1

* Add basin and pre-epic ensemble scripts

* Record the bundled tier, basin, ensemble and MLX Newton runs

* Add the equivalence-margin check on saved ensembles

* Record the seeded keep_best ensemble

* Add the re-measured keep_best figures to ADR 0003

* Name the pre-change code precisely in validation guide

* Regenerate the multi-model ensemble figure

* Rebuild paper.pdf

* Re-state the bundled single-model and multi-model results

* Update keep_best and multi-model figures in AGENTS and guides

* Update the fixture parity figures in AGENTS and progress summary

* Record the bundled tier wall-clock on the M4 Pro

* Update the cross-backend LL claim in the backends guide

* Record the Amari pairing-correction measurement

* Update paper Table 1 bundled rows and the float32 paragraph

* Clarify the benchmark excerpt in the paper

* Rebuild paper.pdf

* Record the gated oracle and full-suite runs

* Add the Phase 15 re-measurement to the changelog

* Record the Newton independent-seed fits

* Record the same-session 70-channel timing check

* Record the unseeded-start check of the issue-145 binary

* Mark the sections of the validation guide that predate the epic

* Pin every reference setting in the Table 1 and run scripts

* Re-state the Newton-enabled basin results

* Give both matching seeds in the paper

* Rebuild paper.pdf

* Document the ensemble's invsigmin difference

* Smooth the ensemble protocol sentence

* Address the docs-claims review findings

* Address the paper review findings

* Give each ensemble distribution its own line style

* Label the external single-model figures as pre-epic

* Reflow the README parity sentence

* Record the external tier run

* Fill the external-tier figures and fix rounding slips

* Rebuild paper.pdf

* Save the pinned parameter files before Phase 17

* Re-check the reference pins after Phase 17

* Pin do_approx_sphere in the reference settings

* Record the test runs after the Phase 17 merge

* Add the issue-351 findings record

* Reword two remaining contrast phrasings in the paper

* Rebuild paper.pdf

---------

Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Filter CI on code paths (docs, paper, paper.pdf, metadata and
non-Python .context records skip; .context Python files and the
tested ensemble.npz still run), and commit the rebuilt paper.pdf
back only on dev and main. Checked with actionlint.
Epic #324: first-class MLX and reference-faithful fitting
auto-bump-dev.yml can push to dev while the PDF builds, which
rejected the plain push after the epic #324 merge. Checked with
actionlint.
* Date the changelog's release headings

Add each release's tag date to its heading, describe the Unreleased
convention in the intro, and record the changelog mechanism.

* Point a root CHANGELOG.md and PyPI at the changelog

* Extract a release's changelog section as release notes

scripts/changelog_section.py prints one release's section with its
relative docs links made absolute, and --check gates a release.
Tested on the real changelog (11 tests).

* Publish release notes from the changelog and check entries

auto-tag.yml uses the release's changelog section as the notes,
with generated PR notes appended. changelog.yml requires an entry
for package-code PRs into dev (skip-changelog label to opt out)
and the release section for PRs into main. Checked with actionlint.

* Record the changelog practice in the rules

* Date releases by their UTC publication date

v0.1.3, v0.2.0 and v0.3.2 carried their tagged commit's date; use
the GitHub release date (or the tag date where no release exists),
and state the convention in the rule.

* Leave code untouched and rewrite reference links in release notes

The extractor skips code spans and fenced blocks, ignores headings
inside fences, and rewrites reference-style link definitions.
Seven edge-case tests (18 total) cover them.
Rename the changelog's Unreleased section to 0.4.0 (2026-09-24)
with a parity summary, and set the version to 0.4.0.dev0 with
sync_version.py so the release chain tags 0.4.0.
@neuromechanist
neuromechanist merged commit 4b67099 into main Sep 24, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant