Repository navigation
Release 0.4.0 - #368
Merged
Merged
Release 0.4.0#368
Conversation
* Add transform and mixing accessors to AMICAMLXNG Ports AMICATorchNG.transform plus get_mixing_matrix/ get_unmixing_matrix/get_sensor_mixing_matrix/get_rho to the MLX backend (issue #287, epic #278 Phase 1). transform derives the unmixing composition from MLX's own _forward rather than transcribing torch's tensor layout: MLX's W is (n_models, n, n), not torch's (n, n, n_models). Fitting is untouched. * Add state_dict/save/load persistence to AMICAMLXNG state_dict/from_state_dict mirror AMICATorchNG's 12-name param set and refusal guards (unfitted, degenerate, non-finite); save/load write a single device- and framework-agnostic .npz (config/extra as JSON-encoded scalars, params as native arrays -- no torch coupling, no pickle). Loading rebuilds the MLX-only per-iteration caches (lgamma table, log|det W|, comp_used) that AMICATorchNG instead recomputes inline on every call. * Add MLX transform tests MLX-only mechanics for issue #287: model_idx validation, unfitted errors, and transform(training X) reproducing the fit's own E-step activations via the transpose identity _forward relies on. Includes the fit-path no-op pin (C7): a short seeded fit is bit-identical to a recorded run from the epic branch tip, proving this phase touched nothing _fit_once calls. * Add MLX transform cross-backend tests Pins transform and the get_mixing_matrix/get_unmixing_matrix/ get_sensor_mixing_matrix/get_rho accessors against a float64 AMICATorchNG twin holding identical fitted parameters, following the existing test_mlx_newton_cross_backend.py pattern. Covers multi-model model_idx routing with nonzero c, and a genuinely rank-reduced fit for get_sensor_mixing_matrix. * Add MLX persistence tests Round trip (state_dict/from_state_dict and .npz save/load): every param bit-identical, config/extra equal, transform output bit-identical pre/post. Sharing and adaptive-switcher state round trip (forced merge, pdftype=1). Refusal guards: unfitted, degenerate, wrong format_version, missing sections, shape drift. * Document MLX transform and persistence support Update the AMICAMLXNG module docstring, docs/api/mlx-backend.md, docs/guides/amica-differences.md's backend table, AGENTS.md's MLX status sentence and the changelog to reflect issue #287 (epic #278 Phase 1): transform/accessors/save-load are no longer "fast-follow" gaps. Remaining gaps are keep_best (Phase 2) and outlier rejection + LLt/MIR (Phase 3). * Persist n_newton_fallbacks in AMICAMLXNG state_dict Plan enumeration for extra omitted this field (an oversight, not a deliberate divergence -- torch's equivalent extra dict includes it); a reloaded model was silently losing the Newton-fallback counter. Round-trip test drives a genuine nonzero count via a deliberately under-determined 256-sample block (same construction as test_mlx_newton.py::test_fallback_ramps_toward_lrate_cap_and_counts) through both state_dict/from_state_dict and .npz save/load. * Harden AMICAMLXNG persistence load path PR review on #305 found the load path under-validated: - _load_params shape-checked only A/comp_list; extend to all 12 _PARAM_ARRAYS against an expectation table read off _initialize_parameters's real allocations. sphere's input-channel width isn't derivable from config under rank reduction, so it's taken from the restored sphere itself and mean is cross-checked against it. - comp_list/pdtype no longer force-cast unconditionally: a float array with NaN or a fractional value is numpy UB once cast to an index dtype. _safe_int_cast requires the value already be integer, or finite and whole, before casting; corrects the "dtype preserved" docstring claim (it force-casts). - load() wraps the npz read so a byte-truncated or otherwise corrupt file (zipfile.BadZipFile/EOFError/OSError/ValueError) re-raises as the same named "malformed AMICAMLXNG save file" error instead of an opaque exception. - extra now requires all 14 keys (format_version 1 has no older files to be lenient for), mirroring the params completeness check, instead of 10 keys raising bare KeyError and 3 silently defaulting to []. - W recomputes log|det W| on load; a finite-but-singular restored W makes slogdet return -inf rather than raise, so that's checked explicitly. - Dropped the dead _INT_PARAM_ARRAYS tuple (only _INT_PARAM_DTYPES is read). Also adds AMICAMLXNG.n_channels_in (port of AMICATorchNG's property), so a reloaded rank-reduced model can still report its original input width, and corrects the stale "this backend has no load path" comment on _sphere_np now that _load_params is a second sphere-assignment point. * Fix stale MLX backend-parity example in rules doc .rules/backend_parity.md's "existing legitimate examples" sentence still said transform raised NotImplementedError and save/load was absent; both landed in epic #278 Phase 1 (#287). Names only the genuinely still-absent items: do_reject, keep_best, LLt/MIR (Phases 2/3). * Add persistence hardening and review-requested test coverage Covers the PR #305 review fixes: - shape-drift rejection for every param (not just A/comp_list), including mean's sphere-derived n_channels_in cross-check - comp_list dtype safety: rejects NaN and fractional values, accepts whole-valued float arrays - singular restored W is rejected (a whole zeroed column reliably reaches slogdet's -inf; a duplicated column on an already float32-rounded W does not -- documented in the test) - a real byte-truncated (not just well-formed-npz-missing-a-key) .npz raises the same named error - get_rho() and state_dict() defense-in-depth: a force-set non-finite rho/param (clean stop_reason) is caught and named - n_channels_in round trips on a rank-reduced fit - rank-reduced save/load round trip: all 12 params + transform bit-identical through both state_dict and .npz - n_restarts=3 round trip: restart_seeds_/restart_lls_/ restart_stop_reasons_ keep their full per-restart length/content, not a vacuous length-1 list - cross-backend accessor test: documents why get_unmixing_matrix alone needs the looser 1e-3 bound (independent per-backend matrix inversion, not a pure reindex/transpose like mixing/rho) * Make fit-path no-op pin survive cross-GPU float32 noise CI failed on a different Apple GPU model: MLX float32 is bit-reproducible on one machine but not across GPU models (observed: ll_history[0] differs from the M4 Pro recording by ~7e-8 relative). Reshape the pin instead of deleting it: - ll_history/final_ll_ now compared per-entry within 5e-6 relative (~50x the observed ~1e-7 cross-GPU spread, below the ~1e-6+ shift a genuine fit-path change produces -- see #216's block_size precedent) instead of exact equality. - Drop the SHA-256 A/W hashes (a cross-machine hash can never match); replace with the same 5e-6 tolerance on five representative A entries. - stop_reason and len(ll_history) stay exact -- machine-independent. - Docstring now says plainly that the bit-level no-op claim was verified same-machine at development time (epic tip 076605c, before/after, bit-identical there); this test is the looser cross-machine canary. * Scale the A-entry no-op spot checks by matrix max, not per-entry CI failed again: A[10, 20] is small (-0.0037), so the per-entry 5e-6 relative bound from the previous fix was an absolute tolerance of ~1.9e-8 there -- tighter than the ~1.1e-7 cross-GPU float32 noise CI actually observed. Same near-zero-entry failure mode _max_rel_disagreement's docstring in test_mlx_transform_cross_backend.py already explains for transform's output; apply the same fix here. The five A spot entries are now compared against a recorded max(|A|) (~1.0, A columns are ~unit-normalized) instead of against each entry's own magnitude. ll_history/final_ll_ keep the per-entry relative check (safe: their magnitude, ~3.3, stays well clear of zero).
test_run_fortran_rejects_sub_resolution_config bet that 8ch x 2000 samples always stamps 0.00 s per iteration; a loaded weekly-run macOS runner crossed one 0.01 stamp and the guard rightly had nothing to reject (issue #308, first weekly failure). The test now asserts the guard's actual contract -- either the named sub-resolution error, or a mean at least one stamp over the timed iterations, never a bogus 0.0 -- so both legitimate outcomes pass. Verified against the real bundled amica15mac binary locally.
* Add keep_best best-iterate safeguard to MLX backend Ports #51's snapshot/restore mechanism from AMICATorchNG onto AMICAMLXNG: fit() tracks the highest-LL iterate and restores it when the run ends more than _KEEP_BEST_TOL below that peak. Inactive under share_comps (a merge changes the parameter count, so an earlier snapshot cannot be reverted to without undoing it). Persisted additively in state_dict's config; ll_history is never rewritten. Epic #278 Phase 2, issue #288. * Add keep_best tests for MLX backend Covers the monotone no-op pin, a genuine forced-restore overshoot (measured, not assumed), ll_history immutability, bit-identical truncated-refit comparison, n_kurt_done/pdtype consistency under pdftype=1, the share_comps disable, and state_dict persistence (including additive-compat for a payload missing the key). All bit-identity checks run within a single test process to stay safe against MLX's cross-GPU float32 non-reproducibility. A cross-backend companion asserts the shared restore CONTRACT (final_ll_ == max(ll_history) when active) without asserting decision equality between PyTorch's float64 and MLX's float32 trajectories. Epic #278 Phase 2, issue #288. * Document MLX keep_best in the backend-parity docs Backend table, gaps sentences (AGENTS.md, mlx-backend.md, backend_parity.md) and changelog now reflect keep_best as implemented on MLX; the remaining MLX gap is do_reject + LLt/MIR (Phase 3). Epic #278 Phase 2, issue #288. * Test the degenerate-stop guard and restart record PR #310 review additions: the restore guard's refusal to rescue a diverged fit that peaked earlier is now pinned in both the MLX and torch backends via the sanctioned error-injection subclass (real iterations, then a forced nan_ll, asserting no restore and bit- identity with keep_best=False); and the keep_best x n_restarts composition is pinned (restart_lls_ reports the restored iterate, bit-identical to a standalone fit's final_ll_). The grad_norm-stop recipe variant is deliberately skipped: the guard is a set-membership test on stop_reason, so nan_ll alone covers the branch. Tested: 13 keepbest tests + torch twin pass; ruff and ty clean.
* Add LLt stash to the MLX backend Port issue #157's per-block E-step stash (AMICATorchNG) so the LLt export never needs a second forward pass. _get_block_updates emits logV/ll_samples per block, _accumulate_blocks(stash_llt=True) scatters them into _llt_logv/_llt_ll as part of the same lazy graph, and _fit_once materializes them into _llt_lht/_llt_lt after any keep_best restore (which now rolls the stash back with the params). Fixes three test doubles that override _accumulate_blocks to accept the new stash_llt kwarg. * Add do_reject outlier rejection to the MLX backend Port issue #123's AMICATorchNG good_idx mechanism: constructor gains do_reject/rejsig/rejstart/rejint/maxrej with torch's names, defaults and validation; _fit_once shrinks good_idx on the reject schedule and restricts the E-step to it; keep_best's track_best now also excludes do_reject. Design decision: the rejection statistic is read FROM the LLt stash (_llt_ll indexed by good_idx) instead of a second forward pass, the NumPy backend's design -- pre-empting AMICATorchNG's open follow-up (#298) to drop its own extra _sample_ll pass. numrej/good_idx persist additively in state_dict's extra (extra.get fallback), the one place phase-1's require-all-extra-keys check is relaxed, documented at both call sites. * Add model_loglik/model_probability to the MLX backend Port issue #141's per-model posterior accessors (AMICATorchNG core.py:3041-3132): score arbitrary data through the stored sphere/mean/W via the existing block-wise _forward, returning (n_models, n_samples) Lht, plus the column-softmax model_probability. * Add write_amica_output EEGLAB export to the MLX backend Thin adapter over numpy_impl.load.write_amicaout, matching AMICATorchNG.write_amica_output (issue #92): converts mx arrays to numpy, transposes W from MLX's (n_models, n, n) to write_amicaout's (nw, nw, num_models) contract, applies the keep_best LL-truncation so the written LL trajectory ends at the returned iterate, and passes Lht/Lt from the materialized LLt stash (warn-and-omit for a from_state_dict/load model). Verified against a float64 torch twin built from identical fitted parameters: W/A/LL/LLt/comp_list agree to float32 precision (~3e-7) or exactly. * Add mir/pmi diagnostics to the MLX backend Port issue #137's MIR/PMI accessors (AMICATorchNG core.py:2929-3036) delegating to pamica.metrics: mir() composes get_unmixing_matrix(model_idx) @ sphere, pmi() runs pairwise_mi on transform(). Ports the #300 fitted-geometry PCA guard (sphere.shape[0] != sphere.shape[1]) rather than a config-only check, since this backend has no explicit pcakeep/pcadb parameter. fit(X, mir_step=...) appends (iteration, mir_nats, variance) waypoints to mir_history_ on the given schedule, excluded from the keep_best snapshot and from state_dict (a true trajectory, not a fitted parameter). * Add tests for MLX reject/LLt/export/scoring/mir/no-op Six new mlx_tests files covering the phase's five deliverables plus the fit-path no-op check, and the one new cross-backend file .rules/backend_parity.md calls for: - test_mlx_llt_stash.py: the Lt.sum()/(n_good*nw) == final_ll_ invariant, keep_best stash rollback, LLt omission with no E-step. - test_mlx_reject.py: good_idx schedule, maxrej cap, none/all-rejected edges, keep_best inactive warning, state_dict/save-load round trip and additive-fallback loading of pre-Phase-3 payloads. - test_mlx_export.py: write_amica_output -> loadmodout round trip, do_reject zero-sentinel, from_state_dict LLt omission, no extra forward pass, and a byte-level comparison against a float64 torch twin's export. - test_mlx_scoring.py: model_loglik/model_probability against the E-step's own forward pass and the one-M-step-behind stash. - test_mlx_mir.py: mir/pmi composition, mir_step schedule, the failed- waypoint NaN guard, the #300 PCA-reduction guard on a real rank-reduced fit, and mir_history_'s keep_best/state_dict exclusions. - test_mlx_fit_noop.py: a default fit (do_reject off, mir_step 0) is bit-identical, same process, to the epic tip before this phase (loaded live via git show as a sibling module). - test_mlx_reject_cross_backend.py: MLX and PyTorch reject the SAME sample set from identical config (evidence for the stash-based rejection-statistic design decision), plus mir() numeric agreement against a torch twin. * Document epic #278 Phase 3: MLX gaps closed Update the MLX backend module/package docstrings, docs/api/mlx-backend.md, AGENTS.md, docs/guides/amica-differences.md (backend table, do_reject/LLt/ MIR sections, a new #298 rejection-statistic design-decision section) and .rules/backend_parity.md to reflect that outlier rejection, the LLt stash, the EEGLAB export, and MIR/PMI are now implemented on MLX, closing epic #278. Adds the Unreleased changelog entry. * Fix test_mlx_fit_noop.py for shallow CI checkouts git show 2e04006 fails with "bad object" in CI's depth-1 checkout, erroring all three no-op tests. Fetch the one commit needed first (tolerating failure), and skip loudly naming the shallow-clone cause if the object is still unreachable, instead of letting git show raise. * Fix _choose_pdfs to use the do_reject good set on MLX _choose_pdfs(X_t) passed the full sphered dataset instead of X_use (the do_reject-restricted good set AMICATorchNG passes), so under pdftype=1 + do_reject the kurtosis-based family decision silently saw the outliers do_reject had already excluded. Fixed to pass X_use. Regression test in test_mlx_reject_cross_backend.py: a pdftype=1, do_reject=True fit's per-source pdtype decision now matches a float64 torch twin bit-for-bit (confirmed to mismatch 1/32 before the fix, 0/32 after). * Guard write_amica_output against degenerate/non-finite models write_amica_output had no refusal guard on either backend: a degenerate (non-finite-LL) or force-corrupted model could be written silently. state_dict() already refuses via a two-layer guard (degenerate stop_reason, then a defense-in-depth isfinite sweep over the param arrays); add the identical guard to AMICATorchNG.write_amica_output and AMICAMLXNG.write_amica_output. The AMICA wrapper's usability gate only protects callers going through it, so a caller using either backend class directly had no gate at all -- and neither backend has one now missing. This also neutralizes the stale-NaN-stash path: a failed final iteration's LLt stash (from before a nan_params/nan_ll break) can no longer reach disk, since the write is refused before the stash is ever read. Tests: direct-marker degenerate refusal + force-corrupted-param defense-in-depth, one pair per backend (torch: test_ng_backend.py; MLX: test_mlx_export.py, NaN-injection recipe matching test_mlx_persistence.py's existing state_dict test). * Pin do_reject through best-of-N restarts and share_comps Two interaction regression tests (PR #311 review), both verified to hold before being pinned: - do_reject x n_restarts: with the winning seed forced to NOT be the last restart (SEEDS=[54,50,42], test_mlx_restarts.py's existing winner-first fixture), good_idx/numrej/A match a solo fit from the winning seed bit-for-bit -- the restore path (_apply_restart_state) is what a last-restart winner would leave untested. - do_reject x share_comps: comp_thresh=0.9 (test_mlx_sharing.py's real-data recipe for a genuine, non-forced merge) combined with a do_reject schedule that also fires in the same fit. Asserts finite params, the merged comp_list/comp_used surviving, and good_idx shrunk alongside it. * Fix mir_step docstring and document the auto-rank-reduction parity PR #311 review: no behavior change. The mir_step docstring's "cannot be known before _preprocess" framing was misleading -- the sphere's shape IS knowable right after _preprocess runs, a few lines later. The real reason there is no upfront gate for automatic mineig/mineig_rel reduction is torch-parity: AMICATorchNG's own upfront gate only ever covers an explicit pcakeep/pcadb request (issue #300's deliberate scope), never the automatic case, which both backends only catch downstream, per-waypoint, inside mir()'s own #300 geometry guard. Reworded the docstring and added a clarifying section to amica-differences.md. Adds the missing test: fit(mir_step=N) on rank-reduced real data (SVD projection) completes normally, with every waypoint recorded as a warned (iter, NaN, NaN) entry. * Match torch's keep_best-inactive reason precedence exactly MLX checked share_comps before do_reject when reporting why keep_best is inactive; AMICATorchNG checks do_reject first (torch_impl/core.py:2425). When both flags are on the two backends must report the same reason for the same configuration. Swapped to match, and pinned with a test where both are enabled simultaneously. * Extend the torch-twin export byte comparison to n_models=2 PR #311 review: parametrize the single-model-only byte comparison to also cover n_models=2, and add an explicit per-model W diff (reshape the raw file back to write_amicaout's own (nw, nw, num_models) contract and compare model 0 and model 1 slices separately) so a per-model swap or a wrong axis in the (n_models, n, n) -> (n, n, num_models) transpose shows up as a named model index, not just "W differs somewhere" in the aggregate flat-byte check. Confirmed non-vacuous: the two models' W differ by ~0.16 in this fit. * Document the direct-call injection-pattern variant in testing.md PR #311 review: test_mlx_reject.py and test_numpy_reject.py already use a lighter variant of the sanctioned error-injection pattern -- calling _reject_outliers directly on a hand-built ll_vec, rather than a subclass, to reach the all-rejected branch a real fit can never organically trigger. Acknowledge it explicitly alongside the subclass form. * Mention the write_amica_output guard in the changelog Amends the Phase 3 Unreleased entry to note the degenerate/non-finite refusal guard added to write_amica_output on both backends (PR #311 review, item 3).
* Attribute degenerate-fit refusal to AMICA wrapper Row 5 of the differences table credited the raw PyTorch backend with refusing transform/get_*/save on a degenerate fit; that guard is the AMICA wrapper's _check_usable contract. The raw AMICATorchNG backend has no such guard of its own (tracked as issue #306). * Document unmapped Fortran keywords Add a section resolving the page's exhaustiveness contract: points to fortran_params.py's FORTRAN_UNSUPPORTED_KEYS as the enumeration, names filter_length/dft_length/decwindow as dead in the reference itself (alongside the already-documented do_choose_pdfs/Spinv2 dead code), and records the do_rho-vs-pdftype divergence: Fortran can freeze the GG shape while keeping pdftype=0, but every pamica backend derives the freeze from pdftype directly, making that combination unreachable. Links issues #312/#314 for the live lifecycle/diagnostics gaps. * Port variance_order to AMICAMLXNG Adds the EEGLAB back-projected-variance component order (issue #92) to the MLX backend, mirroring AMICATorchNG.variance_order with MLX's model-major W layout ((n_models, n, n)) and the unfitted RuntimeError. Closes the one accessor gap the epic's feature-parity audit found in an otherwise-complete Phase 3; updates the module docstring's ported-features claim to match. Adds cross-backend parity coverage against a float64 AMICATorchNG twin holding identical fitted parameters (order matches exactly on a 30-iteration multi-model fit with non-degenerate variance gaps, with a gap-size assertion documenting the tie risk), plus MLX-only unfitted/invalid-model_idx/permutation tests. * Correct validation harness backend coverage claim AGENTS.md claimed the harness runs "both implementations (NG + NumPy)"; validate_implementations.py only exercises torch vs Fortran. NumPy parity lives in pytest (test_sample_data_numpy_vs_fortran); MLX validation lives in mlx_tests/ plus the cross-backend suites. Extending the harness itself is tracked as issue #315. * Fix unsupported-key comment for writestep/do_history The FORTRAN_UNSUPPORTED_KEYS entries for writestep/do_history/ histstep said these have no pamica equivalent; false for the legacy NumPy backend, which implements all three (numpy_impl/core.py's fit loop and _write_history). Keeps the keys unsupported here (this module targets the torch wrapper, which has no matching mechanism) but corrects the reason: NumPy-only, torch/MLX gap tracked as #312. * Refresh stale feature-parity status notes feature_parity.md and progress_summary.md still claimed the runtime- vs-Fortran benchmark was unmeasured and issue #15 (save/load and plot_components test coverage) was open; both landed since. Also corrects progress_summary.md's validation-harness claim (same fix as AGENTS.md: torch-vs-Fortran only, not NG + NumPy). Adds a header note to feature_parity.md pointing to amica-differences.md's table as the live, maintained backend-parity source of truth. * Align NumPy backend's inert pdftype default to 0 self.pdftype is never read after assignment (this backend always runs the GG update; _compute_log_pdf takes no pdftype param), so the old default of 1 was a harmless but confusing mismatch against torch/MLX's constructor default of 0. Aligned for surface consistency only; no behavior change. * Add changelog entry for epic polish round Summarizes the audit-driven fixes ahead of merge to dev: MLX variance_order, the doc corrections, and the new tracking issues (#312, #314, #315) plus the #306 extension. * Scope variance_order precision claim to defaults
A RuntimeError from inside _fit_once's iteration body (the #274 condition-number guard, _choose_pdfs's pdtype invariant, and _pinv_sphere's non-finite-sphere guard on MLX; _update_unmixing_ matrices's torch.linalg.inv failure and _pinv_sphere's guard on torch) fires AFTER that iteration's _update_parameters already reassigned A/mu/beta/rho/alpha/gm/c, but BEFORE W/_logdet_W are updated. stop_reason stayed "max_iter", so a caller of the single-restart fit() path (no try/except there) that catches the propagated exception and persists the model anyway would have every state_dict()/write_amica_output() degenerate-fit guard accept it silently. Set self.stop_reason = restarts.ERROR_STOP_REASON immediately before each raise, so the instance itself -- not just the exception -- is left degenerate regardless of whether the caller catches it. The multi-restart path's own except block already did this; kept as redundant-but-harmless so it does not depend on every raise site doing this correctly. Tests (both backends): a sanctioned injection subclass runs a real fit for N genuine iterations, then forces the target guard's raise on a later iteration -- confirmed to fail without the fix (stop_reason stayed "max_iter") and pass with it, and that state_dict/ write_amica_output then refuse the resulting instance. Plus direct-call coverage for the two guards not reachable via a real fit trajectory.
_snapshot_params/_restore_params rolled back W/rho/etc via _PARAM_ARRAYS, but the two MLX-only per-iteration caches (_logdet_W, feeding log|det W| into every logV; _lgamma_table, feeding the GG log-density) were never captured -- so a restore left them at the LAST (discarded) iterate's values, silently corrupting model_loglik/model_probability/mir on the "restored" model even though W/rho themselves rolled back correctly. Added both to the snapshot dict (captured unconditionally, since both always exist once _initialize_parameters has run) rather than rebuilding them post-restore the way _load_params does: it keeps the snapshot a single self-contained point-in-time capture with nothing a future _restore_params caller could forget, and _restore_params's existing generic setattr loop handles them with no extra code at the call site. Both are already in _RESTART_STATE_ATTRS. Regression test (forced-restore recipe, oracle = a state_dict round trip, which independently rebuilds both caches from the restored W/rho): measured max_diff = 0.3804 (100% of 8192 elements, up to ~0.43% relative error) before the fix, 0.0 (bit-exact) after.
mir()'s ValueError (PCA reduction, or metrics.mir's own near-singular -unmixing check) reflects the fit's geometry -- a fact that does not spontaneously resolve mid-fit -- so every remaining scheduled waypoint re-logged the identical warning on a long fit. LinAlgError stays per-waypoint: that is the genuinely transient case (a near-singular W the natural gradient can pass through) the existing comment already describes. After the first ValueError: warn once (now naming that waypoints are disabled), record that one NaN entry, and stop scheduling waypoints for the rest of this fit (no more mir() calls, no more entries) -- both backends. A local flag inside _fit_once, not an instance attribute: it only matters within one fit() call. Updated the two existing tests this changes the contract of (injected-failure and organic-rank-reduction) to assert the new stop-after-first-ValueError behavior on both backends, including that mir()/flaky_mir is never called again after the failure.
_load_params's finiteness check only covered comp_list/pdtype (via _safe_int_cast) -- the other 10 float _PARAM_ARRAYS (A, W, c, mu, alpha, beta, rho, gm, mean, sphere) had no isfinite check at all, so a NaN/inf-poisoned payload would load "successfully" and only surface later as a confusing downstream NaN with no diagnostic, contradicting the same promise the existing W-singularity check already makes for W specifically. Extends the named-ValueError check to every float param. Test: a NaN-poisoned mu payload (representative of the 10) is rejected by from_state_dict with a named error.
max_iter=0 used to run the EM loop zero times and "complete" successfully with stop_reason="max_iter" (not a degenerate marker) and final_ll_=NaN -- an untrained model that every state_dict()/ write_amica_output() degenerate-fit guard then accepted, since none of them check "did an E-step ever actually run", only "did stop_reason end up degenerate". Added a named ValueError at fit entry on both backends, alongside the existing X.ndim/mir_step checks -- both are argument-validation guards raising before any work happens, and the multi-restart path's narrow "except RuntimeError" already lets ValueError propagate on the first restart rather than swallowing it as a degenerate one. The NumPy backend already validates max_iter >= 1 at construction (pre-existing, unmodified); no fix needed there. Tests: max_iter=0 raises the named error and leaves the model untouched (unfitted) on both backends.
_load_params's numrej/good_idx restoration reached raw int()/ np.asarray() calls with no try/except -- a corrupted or hand-edited payload's malformed value there raised an opaque bare TypeError/ValueError instead of the "malformed AMICAMLXNG state: ..." message every other field in this method already uses. Tests: a non-numeric numrej and a good_idx that cannot be converted to an integer index array each raise the named error.
Mirrors test_mlx_reject.py's do_reject analog. mir_history_ is in _RESTART_STATE_ATTRS, so this exercises the same snapshot/restore path with mir_step set so waypoints actually populate (PR #318 review item 7).
__init__ already validated max_iter, but it is a plain public attribute a caller can reassign afterward, so fit() silently ran the EM loop zero times if max_iter was mutated post-construction. torch/mlx take max_iter as a fit() parameter re-validated every call; numpy takes it only at construction, so re-check self.max_iter at fit entry to close the same gap (PR #318 review item 5 clarification: all three backends).
One fit combines n_models=2, share_comps (forced genuine merge), do_reject, n_restarts=2, and pdftype=1 (with keep_best=True passed to confirm the do_reject/share_comps exclusion is logged, not silently ignored), then walks every accessor, state_dict/from_state_dict, .npz save/load, and write_amica_output/loadmodout, asserting consistency at each stage (PR #318 review item 8).
Both use the deterministic two-model forcing recipe from test_mlx_sharing.py (share_start=4, share_iter=8, comp_thresh=0.9): a genuine merge, then write_amica_output/loadmodout round trips, and mir() stays finite and per-model_idx-aware on the merged model (PR #318 review item 9).
_PARAM_NAMES was missing mean/sphere/pdtype, so the module docstring's 'every fitted array is compared bit for bit' claim was not literally true. All three agree bit-for-bit with the pre-Phase-3 epic tip (PR #318 review item 10).
Both backends' no-E-step LLt tests called fit(max_iter=0) to reach a 'never fitted' state, which item 5's max_iter>=1 guard now rejects before that state is reachable. Rewritten to check the same construction-time invariant directly; the write_amica_output-omits-LLt behavior they also exercised is already covered by test_from_state_dict_write_amica_output_omits_llt (torch) and test_from_state_dict_model_warns_and_omits_llt (mlx).
Four places claimed 'epic #278 is complete as of Phase 3' without naming variance_order, which actually landed in the post-Phase-3 polish round (mlx_impl/__init__.py, .rules/backend_parity.md, AGENTS.md); also fixed docs/changelog.md's do_reject/keep_best exclusion note from future to past tense, now that do_reject shipped in Phase 3 (PR #318 review items 11-12).
Several docstring/comment citations into torch_impl/core.py had drifted from their targets (this epic's own accumulated edits, plus this round's exception-consistency/max_iter changes): write_amica_output, _PARAM_TENSORS, the digamma dorho-gate (also pointed at the wrong code entirely -- the NaN-reset path, not the gate itself), and the phase-1 accessor citations (transform, get_mixing_matrix, n_channels_in, get_rho, get_sensor_mixing_matrix, get_unmixing_matrix, _check_model_idx). All recomputed against the current file (PR #318 review item 13).
docs/api/mlx-backend.md listed the Phase 1 accessors but omitted variance_order, which landed in the post-Phase-3 polish round (PR #318 review item 14).
docs/guides/validation.md's table said rejection was 'both backends' (stale since MLX shipped it in Phase 3) and attributed the degenerate-fit refusal on transform/get_*/save to the raw backends, when it is actually the AMICA wrapper's contract (the raw-backend gap is issue #306); write_amica_output is the one exception, which gained its own raw-backend guard directly in this epic (PR #318 review item 15).
Adds the third variant to the sanctioned exception-injection section: wrapping a real method with monkeypatch to count/record calls while still delegating to the real implementation, used by test_write_amica_output_makes_no_extra_forward_pass to prove zero E-step passes (PR #318 review item 16).
Epic: MLX backend non-fitting parity
* Surface MLX feature parity in user-facing docs Post-epic-#278 discoverability pass: the AGENTS.md map annotation still said GG-only; the EEGLAB guide never mentioned the MLX export path; the backends table row and README undersold the now-complete MLX surface. Docs only. * Address docs review: precision claims and naming The export-validation claim now states float32 precision rather than byte comparison (only comp_list is exact), and the test that seeded the overclaim is renamed to match its actual assertion; the backends table row returns to terse form with the parity statement moved to prose that names the wrapper-only from_params_file caveat (#313); MIR is spelled out on first use in the README; the stale GG-only class one-liner in mlx core is corrected. Export suite 12 passed.
* feat: validate every backend in the harness
validate_implementations.py gains --backend {torch,numpy,mlx}, a comma-separated list, or all. Each backend runs the bundled sample with the harness's canonical params (NumPy through its own key map, MLX through AMICAMLXNG) against one shared reference run; an explicit --backend adds a parity summary table with runtime. The default torch-only run prints the same report as before (diffed).
Tested: dispatch and argument-parsing tests (no binary), and the AMICA_RUN_FORTRAN-gated per-backend parity test against the v0.3.3 native binary (3 passed).
* refactor: run the harness MLX row via AMICA
With Phase 4 merged, the torch and MLX runs share one AMICA(backend=...) path, so the MLX row exercises the wrapper a user calls. Numbers unchanged (same summary table as the direct AMICAMLXNG run).
Tested: harness dispatch tests (29 passed); all-backend run against the v0.3.3 native binary.
* docs: use American English spellings
35 British spellings in comments, docstrings and Markdown (behaviour, colour, neighbour, cancelled, licence, optimiser, ...). Proper nouns (Centre de Recherche...) and 'analogue' left as is; no identifiers touched.
Tested: ruff, typos.
* docs: cite torch core by symbol, not line
Seven torch_impl/core.py:<line> citations outside the MLX module pointed at drifted lines; they now name the method (or are dropped where the sentence already names it). Also replaces a stale pamica.py:802-813 citation (a file renamed in #34) in AMICATorchNG._newton_direction. Fortran line citations unchanged.
Tested: ruff.
* docs: MLX getting-started path and harness rows
getting-started gains a complete Apple Silicon (MLX) section (install, fit/transform/save/load, params file, AMICAICA, precision). The validation guide documents running the harness per backend and records the per-backend parity rows with their expected bars. The differences page records do_sphere=False (#328) and NumPy refit warm-start (#312), both verified against the code. AGENTS.md, changelog and the native-backend and testing pages follow.
Tested: mkdocs build --strict (scratch env); anchors checked in the built site.
* test: keep suite output out of the repository
An audit of the non-slow suite found 63 tests (13 files) writing ./output via AMICA_NumPy's default outdir; the slow CLI, end-to-end and pdf-family tests wrote test_output, _ng_e2e_tmp and scratch_amica_parity into the repo. A session-scoped autouse fixture now runs each session (per xdist worker) from a temporary directory, the harness resolves its bundled inputs from its own location, and the slow tests write under tmp_path. Stale .gitignore entries removed.
Tested: full non-slow suite (1335 passed, nothing written to the repo); coverage reports still land in the repo root with and without xdist.
* test: end-to-end workflow on torch and MLX
New mne_tests/test_end_to_end_workflow.py mirrors our own pipeline on the bundled EEG: average reference, AMICAICA with pcakeep = n_channels - 1 on each backend, get_sources vs transform, a single-component exclusion (exact back-projection, residual kept, rank drops by one), EEGLAB export reloaded by loadmodout with the padded sphere, AMICA save/load, an input.param-driven fit, and torch vs MLX Hungarian-matched sources >= 0.999. Not marked slow (about 5 s); MLX tests skip without MLX or an Apple GPU.
Tested: 12 passed locally (torch and MLX); measured tolerances recorded in the module.
* fix: transpose NumPy sensor mixing matrix
AMICA_NumPy.get_sensor_mixing_matrix returned pinv(sphere) @ stored A without the transpose the torch and MLX backends apply (the stored W is the unmixing transposed, issue #24), so its columns were rows of the true mixing matrix: about 10% off the torch maps after five iterations and no inverse of get_weights() @ sphere. New cross-backend test pins torch/NumPy agreement (4.6e-12 full rank, 9.5e-11 at pcakeep=20), MLX agreement at float32 tolerance, and the identity on every backend.
Tested: new test fails without the fix (4 NumPy cases) and passes with it (10 passed); backend-guard and share-comps suites green.
* chore: remove stale 2025 debug output
pytorch_debug_test/out.txt and pytorch_output_test/out.txt were run logs committed in August 2025 by the long-removed basic PyTorch backend; nothing reads them. The only reference, a typos exclusion for the first directory, goes with them.
Tested: typos.
* docs: pin the harness rows to the v0.3.3 binary
States why the rows name an explicit release binary instead of the resolver default (its cached 'latest' is never refreshed) and how to fetch it.
Tested: mkdocs build --strict.
* Rebuild paper.pdf
* fix: NumPy writes files only with an outdir
AMICA_NumPy defaulted outdir to ./output, so every fit wrote out.txt at construction, writestep checkpoints, history and its final results into the caller's working directory. The default is now outdir=None, which writes nothing (as torch and MLX never do unless asked); an explicit outdir (keyword, params file, or the CLI's --outdir, whose own default stays output) writes exactly as before. Behavior change recorded in the changelog and the NumPy API page; the conftest chdir stays as defense in depth.
Tested: new test_numpy_outdir.py (default leaves the working directory empty with every write path on; explicit and params-file outdir write as before; fails on the old code) and a CLI test that a run without --outdir still writes ./output; NumPy suites (107 passed).
* fix: NumPy viz plots the true maps and sources
plot_components drew stored-A columns (rows of the true mixing matrix) as mixing vectors, and it and plot_pdf_fits built activations as stored W times raw data (no transpose, sphere or mean); plot_model_comparison skipped the sphere. One helper, _component_maps_and_sources, now computes exactly what get_sensor_mixing_matrix and transform return, and all three plots use it. load_results reads a rank-reduced fit's zero-padded sphere instead of raising.
Tested: new test_numpy_viz.py compares the helper and the plotted lines/histograms with the live accessors, full rank and pcakeep=20 (10 passed; 8 fail on the old code).
* test: pin the default harness run and summary
Pins the literal default run (no --backend: one validation_report.txt, no parity_summary.md, exit 0), renders format_parity_summary with populated rows from two real runs (torch as the stand-in reference, NumPy through compare_results) and checks every numeric column, and runs the no-MLX exit over both 'torch,mlx' and 'all'.
Tested: 32 passed, 3 skipped (the AMICA_RUN_FORTRAN-gated parity test).
* docs: fix two more British spellings
unlabelled and relabelled, missed by the first sweep's word-boundary match. Test names containing 'honour' (identifiers) are left as is.
Tested: typos, ruff.
* docs: trim the getting-started MLX route
The MLX section duplicated the backends guide's end-to-end and precision sections; it is now install, AMICA(backend="mlx") fit/transform, an AMICAICA one-liner and the float32 note, linking to the guide for the full workflow. The install heading no longer uses a bare GPU abbreviation.
Tested: mkdocs build --strict; linked anchors checked.
* docs: harness status and test working directory
progress_summary.md's harness line now matches AGENTS.md (every backend, #315), and the testing guide explains the conftest session chdir and how a test resolves repository paths.
Tested: mkdocs build --strict.
* docs: one description of reference resolution
default_reference_binary's docstring is now the single full description of the reference-binary resolution order; the --native-engine help summarizes it and points there. The native-backend page runs the harness with uv run.
Tested: --help renders; harness tests (32 passed).
* test: two-iteration fits for the map gate
The NumPy-vs-torch sensor-map gate (1e-9) failed on the macOS arm64 runner at 1.3e-9: the backends start within 1e-13 and EM amplifies the gap about sixfold per iteration, so five iterations was platform-sensitive. Two iterations keep the same gate with a 1e-12 gap while the pre-fix transposed maps still differ by 2.6-3%. Tested locally: module passes.
---------
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Write and read the EEGLAB sphere column-major (#336) * Fix S-order tests to expect column-major (#336) * Add cross-backend sphere export round-trip test (#336) * Document the sphere-order fix (#336) * Add native-binary sphere orientation test * Compare MLX sphere bytes bit-exactly * Fix sphere-order comments in write_amicaout * Warn about old asymmetric-sphere output * Correct scratch-history note on S order * Polish sphere-order wording in docs/tests
* Add seeded native-binary oracle helper Writes a pamica state in the reference's load_* formats, runs the pinned v0.3.3 binary and reads its raw output arrays back, for element-wise oracle tests (issue #333). Used by the tests that follow. * Normalize component rows in doscaling Each model's stored A block is the transpose of the reference's mixing matrix, so a component is a block row. doscaling normalized stored columns, which is not a change of scale; torch, NumPy and MLX now rescale each component row with the matching mu/beta, exactly as amica15.f90:1843-1851 does (issue #333). Tested: invariance, cross-backend and doscaling=False byte-identity tests; seeded native oracle matches A/mu/sbeta to round-off. Tests that encoded the column rule are updated. * Retune overshoot recipes for the new trajectories With components rescaled, the aggressive-Newton recipes run monotone, so keep_best restores never fired (failures or silent skips). Add newtrate=3.0 and longer budgets where needed, move the pdftype=1 switch to kurt_start=6, pick n_mix=1 for the default min_dll stop, and check the keep_best recompute only after a restore. Tested: every retuned test passes and no longer skips. * Retune two MLX cross-backend test configs The variance-order guard rejected the 30-iteration fit (gap 5e-4), so run 50 (gap 2.4e-3). One two-model sample sat on the rejsig=2.0 threshold between float32 and float64; use rejsig=2.5. Tested: both tests pass for one and two models. * Count scalestep from 1 and validate it The reference parses scalestep but rescales every iteration. Keep it as a pamica extension counted from 1 (iterations s, 2s, ...), so the default 1 is the reference, and reject a non-integer or < 1 value at construction on every backend instead of dividing by zero mid-fit. Record it as row 13 of the differences page. Tested: cadence and validation tests on torch, NumPy and MLX. * Document component orientation in ADR 0006 ADR 0006 states the storage convention (components are rows of each stored block), the doscaling fix and the Phase 8 layout change, and amends ADR 0001; the issue #24 notes point to it. The changelog records the behavior change with the oracle numbers. * Make keep_best-under-reject tests non-vacuous Both the MLX and torch tests ran monotone fits, so final_ll_ equal to the last LL proved nothing. Use the overshoot recipe, require that it restores without do_reject and that the do_reject trajectory also ends below an earlier peak, then assert no restore. Tested: both tests pass; the old configs were monotone. * Check reject agreement up to float32 rounding Exact agreement of the rejected set is config luck: float32 and float64 per-sample LLs differ by up to ~0.1 by the second pass. Record each pass, require every disagreement to lie within the band measured on that pass, cap the band and the disagreement count, and add controls where a threshold or trajectory divergence fails. Tested: seeds 7 and 8, one and two models, at rejsig 2.0. * Note ADR 0003 figures predate the row rule * Rebuild paper.pdf * Derive MLX pin tolerances from float32 study * Check variance order against a float32 band * Band-check the pdftype1 reject test * Use semantic line breaks in ADR 0006 * Note Phase 8 aligns the initial A norm * Word pre-fix norm ranges as history * Note parity rows predate the row rule * Document when scalestep is validated * Say oracle maxima span backends and models * Mirror the zero-norm guard in Newton test * Test the rescale on Newton-reached states * Rescale a permuted comp_list across backends * Test zero- and NaN-norm rows stay untouched * Fail, not skip, on recipe preconditions * Never read stale or partial oracle output * Check out full history in pytest CI jobs * Fail byte-identity pins in CI, not skip * Use the CI-proven overshoot recipe in torch --------- Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Count schedule gates from 1 like the reference Route every iteration-schedule gate in the torch, NumPy and MLX backends through a shared pamica/schedule.py. The Newton switch (iter .ge. newt_start), the numdecs reset on the switch-on iteration, the iter > newt_start ceiling ratchets, the outlier-rejection schedule and NumPy's restart-on-NaN window compared the 0-based loop index with the reference's 1-based setting and fired one iteration late. Share, kurtosis, A-freeze and writestep/histstep gates were already 1-based and now use the same helpers. Seeded against the v0.3.3 binary (newt_start=3, doscaling off, 6 iterations) the LL deviation drops from 1.2e-3 to 5.4e-11 (torch) and 2.1e-12 (NumPy). * Update schedule-dependent tests to 1-based gates Tests whose configurations or replays encoded the old 0-based newt_start/rejstart: the aggressive-Newton keep_best recipe moves from newt_start=1 to 2 (the exact pre-change trajectory its documented measurements came from), the LLt reject tests fire on iteration 6 of a 6-iteration fit, and the MLX ratchet/reset replays spell out the reference's 1-based arithmetic. Tested: the 11 affected files pass (179 passed, 1 skipped). * Test schedule gates against 1-based iterations Cross-backend real-fit tests (torch, NumPy, MLX): the first Newton step is iteration newt_start, the rho-rate ceiling ratchet opens only after it, the switch-on counter reset follows the reference's replay, rejection first fires on iteration rejstart, and NumPy's restart-on-NaN window covers the first restartiter iterations. Tested: 15 passed; against the pre-fix backends 10 of them fail. * Add gated native oracle for the Newton start Seeds the pinned v0.3.3 binary with pamica's own init (indir/load_*), newt_start=3, doscaling off, single-threaded, 6 iterations. Torch and NumPy match its LL trajectory and A to round-off (torch LL 5.4e-11, A 4.2e-10; NumPy LL 2.1e-12, A 1.1e-10) where the pre-fix code deviated by 1.25e-3 and 3.1e-2. A no-Newton control shows the compared window is Newton-sensitive. Tested with AMICA_RUN_FORTRAN=1: 3 passed; 2 fail on the pre-fix backends. * Document the 1-based schedule gates Changelog entry for issue #335: which gates moved, the native-oracle numbers before and after, the default-path cases that change, and how to reproduce an old trajectory. No row in docs/guides/amica-differences.md changes: the anchored share A-freeze is already recorded there. * Rebuild paper.pdf * Validate iteration-schedule settings One shared validator, schedule.validate_iteration_setting, called by all three constructors with identical messages: newt_start must be an integer >= 0 whether or not do_newton is on (it also gates the rho-rate ratchet), and rejstart an integer >= 1 when do_reject is on (with 1-based counting, rejstart <= 0 skipped the reference's unconditional first pass; NumPy accepted 0, torch and MLX any value). NumPy also validates restartiter >= 0, maxrestarts >= 0 and histstep >= 1 under do_history (histstep=0 was a bare ZeroDivisionError mid-fit), and documents that restartiter=0 disables restart-on-NaN. schedule.every and periodic_due raise ValueError for a non-positive interval. Docstrings: newton_active cites amica15.f90:1666, newt_start=0 and 1 are stated to fit identically, and the MLX newt_start/newtrate entries document the newtrate ratchet. Tested: the new validation, restart-window and interval tests pass (36 passed with the reject suites). * Reflow MLX schedule docstrings The share_start/share_iter and kurt_start entries no longer break mid-clause. No code change. * Record NumPy restart-on-NaN difference Row 14 of the differences guide: the reference allows maxrestarts+1 restarts and, because startover is never cleared (amica15.f90:1046), never calls update_params again after the first one (:1115-1122); the NumPy backend allows maxrestarts restarts and resumes fitting. PyTorch and MLX have no restart-on-NaN path and stop with nan_ll. * Correct the schedule changelog entry Use the measured pre-fix deviations (1.25e-3 over 6 iterations, 1.33e-3 over 100) in the changelog and the oracle test docstring, list restartiter among the settings that now count from 1, and add the new validation. * Test schedule gates at their boundaries The arithmetic test gains the =1 rows for every helper (newton_switches_on at 1 and 0, past_newton_start at newt_start and newt_start+1, rejection with rejstart=1/rejint=1 through maxrej exhaustion, every/periodic_due at 1, the restart window at 1 and 0). The first-Newton-step test runs at newt_start 1 and 5, and a new real-fit test on every backend shows newt_start=0 and 1 give bit-identical trajectories on an overshooting run whose ratchets fire. Tested: 12 passed. * Make the switch-on reset test discriminating The replay-based reset tests passed with the reset one iteration late too. The cross-backend test now uses a searched configuration (4096 frames, lrate=0.8, newt_ramp=1, maxdecs=3, newt_start=3) whose likelihood decreases on iterations 2, 3 and 4, where the reference rule predicts no ratchet and the one-late rule predicts one; the fitted lrate ceiling must match the reference. It fails on all three backends with only newton_switches_on reverted and on the pre-fix backends. The MLX twin, which could not discriminate, is removed in favor of it. Tested: 3 passed; 3 failed with the reset reverted. * Test the rho-rate gate at newt_start=20 A searched natural-gradient configuration (4096 frames, seed 4, lrate 0.5, maxdecs 2) completes a maxdecs cycle on iteration 21 on all three backends, so the rho-rate ratchet gate is now pinned at the shipped default: open for newt_start=20, closed for 21. This is the one default-path change of issue #335. Tested: 3 passed; 3 fail on the pre-fix backends. * Document the reset-test configuration search States the full search that selected the switch-on reset configuration, including the two-model hit found after the first commit. * Shift Phase 7 recipes to 1-based settings Phase 7 retuned its overshoot recipes under 0-based gating. With newt_start and rejstart counted from 1, n becomes n+1 so each documented trajectory reproduces bit for bit: the min_dll overshoot recipe (newt_start 2, newtrate 3.0) again peaks at 70 and stops at 71 on torch (1.9e-4 below the peak) and at 63/64 on MLX (4.2e-4); its do_reject variants move rejstart 5 to 6; the doscaling-rows NEWTON state uses newt_start=2; the min_dll-reachable test uses 6 and the slow rholrate ratchet test 51. Tested: the affected suites pass. * Route scalestep through the schedule module Phase 7's inline (iteration + 1) % scalestep rule becomes schedule.every in all three backends, and its scalestep check becomes schedule.validate_iteration_setting (same semantics and message; still only when doscaling is on), which drops the numbers imports it needed. Tested: test_doscaling_rows passes. * Re-search the newt_start=20 gate configuration Phase 7's component-row doscaling moved the old configuration's second ratchet off iteration 21, and the test's precondition guard failed as intended. The re-searched configuration (4096 frames, seed 1, lrate 0.6, maxdecs 3, newt_ramp 1) decreases at ll indices 4, 5, 8 and 9, 19, 20 on all three backends (smallest decrease 4.7e-4), ratcheting at 8 and 20. Tested: 3 passed. * Note the combined doscaling and Newton-start parity Changelog: with doscaling on (component rows, issue #333) the seeded 100-iteration Newton comparison stays within 3.9e-6 (torch) and 7.0e-6 (NumPy), where the #333 fix alone left 2.1e-4; scalestep now uses the shared schedule helper. --------- Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Add a loader for pre-change pamica in tests pamica/tests/pre_change.py imports the whole package at a pinned commit from git under a private name, so byte-identity tests can compare the live backends with the pre-change ones in one process. It fetches only when the object is missing (a --depth 1 fetch in a full clone makes it shallow), and fails under CI but skips locally when the commit is unreachable. * Store mixing matrix components as rows All three backends now hold A as (n_comps, n_channels): row k is the reference's column A(:, k), and model h's block is A[comp_list[:, h], :], the same matrix as before, so W = inv(block) is unchanged (issue #334). comp_list now indexes components in A as it does in the densities, so: - the share metric compares pinv(sphere) @ A.T, the components' scalp maps, not stored columns of each block; - a merge ties two components' mixing vectors and densities, and the gm-weighted dAk/zeta step averages one component across models; - doscaling rescales each component row once, shared or not; - accessors, W, the EEGLAB export (reference layout), load_results and the viz helpers follow the new layout. Fits without a merge are byte-identical on every backend; ndtmpsum moves by float round-off (it sums per component row, like the reference). Existing sharing, doscaling and MLX tests are updated to the row layout; two short sharing recipes are retuned because early scans now merge most components. Tested: full non-slow suite and slow tests. * Version saved models for the row layout The PyTorch state_dict is now format_version 4 and the MLX save format 2. A version 3 (PyTorch) or 1 (MLX) payload is converted through the shared pamica.component_layout.rows_from_legacy_columns: without loss when its comp_list is unmerged, and refused with a request to refit when a component id is shared across models, since that merge was made under the old column semantics. AMICA.save files (still version 2) go through the same path. Tested with test_component_rows_persistence.py (24 passed): round trips in the new formats, and payloads written by the pre-change classes loaded from git, converted or refused. * Refuse a stale multi-model A in load_results load_results now checks that each model's block of A inverts the W written beside it (max |A_h W_h - I| <= 1e-3; measured at most 3.8e-7 for float32 MLX exports and 1.1e-15 for float64). A multi-model directory written before issue #334 stored A with components as columns and reads back scrambled (residual 1.5); it is refused with the remedy instead of handing the viz helpers wrong sensor maps. Single-model files are the same bytes in both layouts and still load. * Test component rows and sharing semantics test_component_rows.py pins, on the sample EEG: - byte identity of every unshared configuration against the pre-change package (3627a6e) on PyTorch, NumPy and MLX: 1-3 models, doscaling, Newton, pdftype 0/1 and do_reject; - the metric compares get_sensor_mixing_matrix columns, a planted duplicate merges, and a merged group shares one mixing vector; - the EEGLAB A file is the reference layout, single-model exports are byte-identical, and load_results refuses a pre-change multi-model A; - opt-in (AMICA_RUN_FORTRAN=1): updates from a merged load_comp_list state match amica15mac to float64 round-off on PyTorch and NumPy, with doscaling on and off. All pass locally, including the gated oracle. * Document the component-row layout ADR 0007 records the row layout, the sharing semantics that follow from it, the persistence and export policy, and the measured evidence; ADR 0006 and 0001 are amended, and ADR 0006's initial-A normalization line now points to issue #341. The changelog, the differences page (sharing rows, a new section for issue #334, and row 15: the reference's single-precision density normalizers), and the backends, EEGLAB, validation and API pages describe the new behavior. mkdocs build --strict passes. * Reflow two sharing test docstrings Docstring text only; no test logic changes. * Skip MLX wrapper routes without MLX The merged-save refusal test's wrapper route fitted an MLX model without first checking MLX was installed, so it failed on the Linux runner. _wrapper_fit now performs the check for every wrapper route. Tested with MLX hidden: the new test files pass or skip. * Pin byte identity to the Phase 9 epic head The pre-change package is now 0930c0e, the epic head with Phase 9's schedule gates and without this phase, so the two sides differ by the layout change alone. The full byte-identity matrix (Newton included) and the persistence tests pass against it: 92 passed, 2 gated oracle tests skipped. * Stop tests from fetching pinned commits Tests that load historical code ran `git fetch origin <sha> --depth 1`, which in a full clone records a shallow boundary that gc can prune behind (issue #343). pre_change.py now only reads: it checks the pin with `git cat-file -e` and, when the commit is missing, fails under CI and otherwise skips with the command that fetches it (`git fetch origin <sha>`, or `git fetch --unshallow origin` in a shallow clone). Pins must be full SHAs. test_mlx_fit_noop.py, the last loader that fetched, now imports the whole package at its pin through the same helper; its fits are still bit-identical. test_pre_change_loader.py runs the loader behind a git wrapper that logs its arguments and execs the real git, and asserts no fetch ran and the shallow state is unchanged (a negative control with a real fetch trips the assertion). Tested: loader and fit-noop tests pass, also with CI set. * Record the single-precision density constants The reference writes its density normalizers as dble(<literal>), which gfortran reads in single precision: log(dble(1.772453851)) in the rho == 2 branch is 3.0e-8 above log(sqrt(pi)) (measured in the seeded oracle), and the Gaussian and cosh families' normalizers differ from pamica's double values by 3.7e-10, 2.0e-8 and -2.1e-8 (computed). Row 16 of the differences page records this, citing issue #344, which decides whether to adopt the reference's values. The torch and MLX comments that called _LOG_SQRT_2PI and _LOG_NORM_COSH_* bit-for-bit with the binary are corrected, and the merged-state oracle's maxrho=1.99 cites the issue. Comments and docs only. * Show early mass merges happen in the reference With share_comps on, the pinned binary's scan runs but never merges: Spinv2 is never allocated, so every similarity is NaN, even at comp_thresh=0. Its similarity (Spinv2 = Spinv^T Spinv) applied to its own state after 8 iterations from pamica's initialization merges exactly the pairs the torch and NumPy scans merge (32, 32 and 30 at 0.9, 0.95 and 0.99). In the recipe that went non-finite, the reference's update from pamica's post-scan states collapses in step: the second model's gm falls to 5.04e-3 on both after the first scan, and after the second reaches zero in one iteration and NaN in the next. Two gated tests pin this (AMICA_RUN_FORTRAN=1, both pass). The differences page, ADR 0007, the changelog, the validation table and the torch/MLX comments that called the scan unrunnable now say what it does. * Harden the pre-change loader and its guard pre_change.py now fails under CI, and otherwise skips, with its own message when git is not installed (instead of a raw FileNotFoundError) and when the tests do not run from a git checkout (instead of advising a fetch that cannot help). The repository root is a parameter, and a cached alias must map to the same commit or loading raises. The test's PATH wrapper now refuses fetch, pull and every depth or shallow option without running git, so the committed negative control (a fetch sent through _git) is caught by _assert_no_fetch and leaves the clone untouched. New tests cover missing git, a non-checkout root and alias reuse. Tested: 10 passed, also with CI set; the clone is not shallow afterward. * Validate comp_list dtype in the layout conversion rows_from_legacy_columns now refuses a non-integer comp_list (float or NaN ids) with its own ValueError before the bounds check, instead of letting it reach the indexing. test_component_layout.py adds direct unit tests on small arrays for every ValueError branch: the comp_list rank, the dtype, the A shape, an out-of-range id, an id repeated within a model, and an id shared across models (the refit message), plus a block-for-block conversion of an unmerged, permuted comp_list. Tested: 11 new tests pass; persistence tests unchanged. * Refuse a truncated A in load_results load_results now checks that the A file holds num_comps * nw values before reshaping it to (num_comps, nw), and names the file and both counts when it does not. Before, a length no reshape accepts surfaced as a bare NumPy reshape error, and one short by a whole component reshaped to the wrong width and failed later as a matrix-product error. Tested on truncated copies of a real two-model export (both cases). * Raise instead of assert in NumPy _write_results The fitted-A and outdir preconditions of AMICA_NumPy._write_results are now explicit RuntimeErrors, since assert is stripped under python -O. The comment above them also said the reference omits A from its output; the binary does write it, in the layout pamica now writes. Tested: NumPy write and load tests pass. * Widen the byte-identity matrix The unshared byte-identity test now also covers the Gaussian, logistic and sub-Gaussian cosh families (pdftype 2, 3 and 4, single model, n_mix 1 for family 4), four blocks of 1024 samples, a keep_best fit whose log-likelihood falls so the safeguard restores an earlier snapshot (asserted), and best-of-two restarts, on every backend that supports each. The NumPy backend has only the generalized Gaussian and no keep_best, so those cases run on PyTorch and MLX; the test says so and test_numpy_rejects_other_pdftypes pins the first reason. The helper now lets a configuration override its defaults, and compares final_ll_. Tested with CI=1: 66 passed. * Check the MLX double conversion on save again test_old_mlx_save_converts_losslessly now saves the converted model again, checks the file is format 2 with a component-row A, reloads it, and compares A and every output with the saved model, as the torch test does: a second load must not convert the already-converted A again. Tested: both recipes pass. * Fix the Spinv citation and sharing comments The torch _identify_shared_comps docstring cited Spinv's allocation at :551; it is amica15.f90:569. A test_mlx_newton docstring said merged-away columns and curvature not indexed by mixing column; with components as rows the words are component and mixing vector. Comments only. * Label the merged-oracle figures consistently ADR 0007 and the changelog now say the 3-iteration oracle figures are the worst of doscaling on and off, and give the old column semantics' error as the matching worst value (A 0.21, LL 4.3e-4, not 0.20 and 4e-4); the test comment lists both measurements. Docs and comments only. * Describe component-row sharing in the notes AGENTS.md and .context/progress_summary.md described the old column metric and merge as working. They now describe the component-row semantics (ADR 0007, #334): the metric compares de-sphered component mixing vectors, the merge ties components, and the reference binary's own scan never merges because Spinv2 is never allocated. Docs only.
* Reject unknown AMICA_NumPy keyword arguments AMICA_NumPy(**kwargs) forwarded every keyword into a params dict read with params.get(...), so a typo (max_iters=50) or an option the legacy backend does not implement (keep_best=True) constructed silently and had no effect. __init__ now validates **kwargs before any other processing: an unrecognized name raises TypeError with a difflib-based suggestion, and a name implemented on the PyTorch backend (AMICATorchNG) but not this one raises TypeError naming it and pointing to AMICA(backend='torch'). The accepted-keyword set is derived from _CONSUMED_KEYS/_CANONICAL_TO_NUMPY_KEY (the same source the params-file routing uses) plus this constructor's own named parameters, so the two surfaces cannot drift apart. The canonical spelling of the three renamed settings (min_nd/maxdecs/share_iter) is now also accepted as a keyword argument directly, translated the same way a params file's setting already is; passing both spellings at once raises TypeError. * Add tests for AMICA_NumPy kwargs validation Real construction only, no mocks (no .fit() needed to exercise __init__'s validation): a typo raises TypeError with a suggestion; an unsupported cross-backend option raises the specific message; the three canonical-spelling aliases are accepted and their conflict is rejected; every documented NumPy option constructs, alone and all together; and a source scan proves _CONSUMED_KEYS matches every params.get(...)/raw_params.get(...) read site exactly, so the accepted-kwargs set cannot silently drift from what __init__ actually reads. * Document AMICA_NumPy kwargs rejection behavior change docs/changelog.md: Unreleased entry noting the legacy NumPy backend behavior change (unknown or unsupported options now raise instead of being ignored). docs/api/numpy-backend.md: a short paragraph mirroring the existing pdftype callout. * Improve AMICA_NumPy kwarg validation (review) PR #347 review follow-up, several items together since they touch the same functions: - files/data_dim/field_dim now work as constructor keywords, with the same meaning as in params_file, instead of passing the accepted- kwargs check but staying silently inert (fit() with no data still said "no 'files' configured" even after a caller set one via a kwarg). Conflicting values between params_file and a kwarg raise; the same value from both does not. - n_models/n_mix (AMICATorchNG's spelling of this backend's own num_models/num_mix) get a message naming the correct spelling instead of the generic (and here wrong) torch-only message. - n_channels gets its own message: it is inferred from the data passed to fit() on every backend, not a constructor keyword on any of them (the AMICA wrapper's own fit(**kwargs) excludes it too). - Every offending keyword in one call is named in a single error, however many different categories are mixed together. A first version raised on the first category found, so AMICA_NumPy(n_models=2, totally_bogus=1) named only n_models and dropped totally_bogus. - _torch_only_options() imports AMICATorchNG lazily, inside its own body, called only on the error path, so numpy_impl carries no module-level dependency on torch_impl. - Groups difflib/inspect with the other stdlib imports. - Hardens the source-scan test: fails loudly on a params[...] bracket read or a non-literal params.get(...) key, either of which would be invisible to the scan. - Adds typo-suggestion tests for params_file/use_tqdm alongside verbose. * Fix dropped keyword in AMICA.fit mlx check PR #347 review item 4: the same single-category drop AMICA_NumPy._reject_unknown_kwargs had (issue #346) existed here too -- backend='mlx' with fit(dtype=..., blocksize=...) named only 'dtype' (ValueError) and silently dropped 'blocksize' from the message, since the torch-only check raised before the unknown-keyword check ever ran. Both categories are computed before either is raised; a combination now names both in one TypeError. The two single-category messages are unchanged (still ValueError for torch-only alone, TypeError for unknown alone), so existing callers/tests keep working. * Document NumPy kwargs review follow-up docs/changelog.md: a Review follow-up sub-bullet under the Phase 14 entry. docs/api/numpy-backend.md: mentions files/data_dim/field_dim as keyword arguments and the n_models/n_mix/n_channels-specific messages.
* Split the M-step into direction and update Compute the natural-gradient/Newton step and its norm (the reference's dAk and ndtmpsum) in _update_direction and apply it in _update_parameters, in all three backends, sharing one set of intermediates. The fit loops call both back to back, so every fit is byte-identical (checked on 8 configurations per backend). * Follow the reference's iteration order Run the likelihood-decrease response and the stopping checks before the parameter update, and exit on a stop before any parameter moves, as amica15.f90:1015-1122 does, in all three backends. The step taken right after a decrease now uses the halved rate, and a stopped fit returns the parameters whose likelihood is ll_history[-1]. Keep the reference's working rho rate and its ceiling apart (rholrate, rholrate_cap), reset inside the A-update branch. Fits without decreases or stops are byte-identical; seeded oracle deviations drop from 1.4e-2 to 2.7e-5. * Hold A on the reference's schedule without sharing Apply the reference's A-update guard (iter >= share_start and mod(iter, share_iter) <= 5, amica15.f90:1803) in every fit, share_comps on or off, through schedule.share_freeze, replacing the anchored post-merge window. Every backend now rejects share_iter < 7 (NumPy: share_int) whether or not sharing is on, since a shorter cycle would never update A again. Default fits of 100 or more iterations change. * Add iteration-order tests and native oracle Always-on cross-backend tests for the freeze iterations, the decrease timing, stop semantics, keep_best pairing and share_iter validation, plus an opt-in comparison against the native binary (AMICA_RUN_FORTRAN=1). All pass; the native comparison fails on the pre-change code. * Retune tests to the reference's iteration order Existing tests that pinned the old order (update before the checks, a freeze gated on share_comps and anchored on share_start, one rho rate) now pin the reference's, or use recipes that still exercise what they test. Full non-slow suite and the gated native oracles pass. * Document the reference's iteration order ADR 0008, changelog entry with oracle and parity numbers, and the differences guide: the A-freeze is now the reference's arithmetic, and convergence stops return the parameters final_ll_ describes. mkdocs build --strict passes. * Test the rho-rate ceiling as its own value Rate tests that read rholrate as the ceiling now read rholrate_cap, the cross-backend state copies carry it, and a new test round-trips it through state_dict on PyTorch and MLX, including a save without the key. Affected tests pass, including the slow torch ratchet test. * Stop overshoot recipes on the first decrease The keep_best overshoot recipes stopped via min_dll=1e-4/maxincs=2, which under the reference order ended at the peak or below it depending on round-off: at the peak on the macOS CI runner, and in 1 to 2 of 12 data perturbations locally. maxincs=0 with min_dll=1e-8 stops them on their first likelihood decrease instead, which held in all 12 perturbations on PyTorch and MLX, with and without rejection. Affected tests pass. * Stop identically on non-finite values Every backend now keeps a non-finite likelihood out of its history (NumPy recorded it), stops before applying a non-finite step or gradient norm (new degenerate stop_reason nan_direction), and stops on non-finite parameters right after an update (nan_params, now also in PyTorch and NumPy). NumPy defers only an A/W corruption inside its restart window to restart-on-NaN, which can repair it. Cross-backend injection tests pass. * Clear numincs on a NumPy restart The reference's min_dll comparison with the NaN likelihood of its restart iteration is false, so it zeroes numincs (numdecs is left as it was). NumPy's restart carried the small-gain count across; it now clears it. The new test stops at index 8 as expected and at 6 without the fix. * Validate saved learning rates on load from_state_dict (and the MLX file load) now refuses a missing or non-finite lrate, lrate_cap, newtrate, rholrate or rholrate_cap with a ValueError naming the field; a missing rholrate no longer surfaces as a bare KeyError on PyTorch. Tested on PyTorch and MLX. * Validate share_start on every fit The A-freeze reads share_start whether or not share_comps is on, so every backend now requires an integer share_start >= 1 always. The share_iter message now says why a value below 7 is refused (A held permanently from share_start), and NumPy's names both of its spellings. Both checks are reached through a Fortran params file on all backends; tests pass. * Test every convergence stop's parameters The stop test now runs min_dll, grad_norm, grad_norm_floor and lrate_floor on every backend, each with a precondition on its stop_reason and a twin that runs the same iterations to max_iter. All 12 cases pass. * Test that a NumPy restart skips the update A pass-through recorder on _update_parameters shows the restarting iteration applies no update; the same fit of the pre-change code, loaded through pre_change.py as a control, updated on it. Passes. * Test MNE n_iter_ after each kind of stop to_mne_ica().n_iter_ equals iteration + 1 and len(ll_history) after a min_dll stop and after a max_iter fit, on PyTorch and MLX. Passes. * Take the oracle floor over two thread pairs The reference noise floor is now the larger deviation of a 2- or 4-thread binary run from the 1-thread run. The multi-threaded runs are not reproducible, so the docstring gives each floor's range over three runs. All four gated oracle cases pass in each run. * Fix docstrings and Fortran citations mir_history_ records no waypoint on a stopping iteration; MLX documents its rholrate/rholrate_cap pair as PyTorch does; NumPy fit() gets the iteration-order notes. Rejection is cited as amica15.f90:1136-1140 and the check ranges start at 1056 and 1019. Docstring and comment changes only. * Document the non-finite stops and validation ADR 0008 states the final_ll_ invariant with its keep_best qualifier and the review's changes; the changelog and differences guide cover the non-finite stops, share_start validation and saved-rate checks, with the two-pair oracle floors. AGENTS.md and the context notes describe the unconditional A-freeze. mkdocs build --strict passes.
* Use the reference's single-precision constants The reference writes its density normalizers as default-kind literals widened with dble, so the binary uses their float32 roundings. All three backends now read those values, and epsdble's, from one shared module, pamica/reference_constants.py; MLX holds the rho == 2 normalizer in its lgamma table. The literal-formula tests and the benchmark's copy now use the reference's values. Tested: pdf-family and backend test files pass. * Compare pre-change fits at the old constants The byte-identity tests against code from before issue #344 (component rows, doscaling off, the MLX pre-Phase-3 tip) now give the live backend its old density constants through pre_change.use_pre_344_constants, so they still isolate the change each was written for: their two-model and non-GG fits reach the changed normalizers. Tested: the three files pass. * Pin the constants against the reference test_reference_constants.py checks each shared constant against the float32 rounding of its literal on the cited reference line (exact rational arithmetic), pins the sweep of amica15.f90 and its header, checks that no backend keeps a copy, checks every backend's rho == 2 log-density, and (opt-in) seeds the native binary for pdftype 2, 4 and 1. Tested: 23 passed with AMICA_RUN_FORTRAN=1. * Run the merged oracle at rho == 2 The merged-state oracle no longer holds maxrho at 1.99: its warm seed now has mixtures clamped at 2, it asserts that, and the code before issue #344 is replayed from the same state as a control (off by 2.8e-9 in log-likelihood after 1 iteration, against 2.2e-15 now). Tolerances unchanged. Tested with AMICA_RUN_FORTRAN=1. * Document the single-precision constants Changelog entry with the oracle numbers; the differences guide gains a section on the constants, and row 16 now records the one kind of single-precision literal pamica keeps at its decimal value (defaults of input.param keys); the validation guide notes the new family oracle. Tested: mkdocs build --strict, typos. * Note the reference's unset lrate0 in row 16 Without an lrate key the reference never sets its ceiling lrate0, so its lrate drops to 0 after the first iteration (checked against the pinned binary). Row 16 now says so and uses comp_thresh as the example. Tested: mkdocs build --strict, typos. * Skip MLX before patching its constants On a host without MLX the byte-identity tests imported the MLX module to patch it before the check that skips them. Resolve the classes first. Tested: the two files pass with MLX and skip its cases without. * Remove the unused compute_log_pdf numpy_impl.pdf.compute_log_pdf had no callers in the package, tests or docs. Tested: the pdf and viz tests pass. * Plot the fit's density in compute_pdf compute_pdf, which viz.plot_pdf_fits draws, now takes its normalizers from pamica.reference_constants, including the reference's single-precision sqrt(pi) at rho == 2, so it draws the density the fit uses. Its docstring says it is a plotting helper with the legacy pdftype numbering. Tested: the pdf, viz and constants tests pass. * Scan every backend module for copies The no-copy test now scans every .py file of torch_impl, numpy_impl and mlx_impl with comments and strings blanked, and flags the linear forms (sqrt(pi), sqrt(2*pi), ...) as well as the log forms and the literals. A new test checks that it flags every definition this change removed. Tested: passes; fails on the old pdf.py sqrt(pi) line (checked by hand). * Assert the family held in the oracle The family oracle now asserts that every source kept its family and no kurtosis switch ran (n_kurt_done == 0), for pdftype=1 with num_kurt=0 too, in the live and the pre-change runs. Tested with AMICA_RUN_FORTRAN=1. * Check the pre-#344 constant substitution A control test: every constant use_pre_344_constants patches differs from the live value, the patch takes effect, it covers every changed constant each backend module reads, and on MLX it restores the pre-change rho == 2 table entry. Tested: passes on all three backends. * Note the plotting helper in the changelog compute_pdf now draws the fit's density and compute_log_pdf is gone. Tested: mkdocs build --strict, typos.
* Run the native binary from its own draw run_drawn_reference runs the binary with nothing loaded and a fixed seed; with max_iter=0 it writes its drawn initialization. The run and read-back are shared with run_seeded_reference. * Normalize the drawn initial mixing matrix Every backend now draws its initial A through the shared pamica.initialization.initial_mixing: per model, 0.01*(0.5-u) with the diagonal set to one and each component divided by its norm, as the reference does (amica15.f90:805-823), including the NumPy restart redraw. A supplied or loaded A is used as is. Tested in later commits. * Start pre-change classes from the new init The byte-identity pins (fb13d76, 0930c0e, 2e04006) and the #339 oracle control now fit the historical class from the live normalized initial A through pre_change.with_normalized_initial_mixing, so each still isolates its own change. Tested: all those byte-identity tests pass. * Test the normalized initial mixing matrix Always-on cross-backend start checks with a pre-change control, the supplied/loaded A paths, the NumPy restart redraw, and two gated native oracles. Tested: 34 passed with AMICA_RUN_FORTRAN=1. * Re-search the MLX pdftype=1 restore recipe The normalized initial A made seed 3 at lrate 0.5 climb monotonically to max_iter, so the restore was never exercised. Seed 6 at lrate 0.4 (kurt_start 7) overshoots by 0.17 after all five switch passes, on all 12 perturbed inputs. Tested: test_mlx_keepbest.py passes. * Re-record the MLX fit-path canary The normalized initial A moved the recorded trajectory (ll_history[0] by 1.55e-5). Pins re-recorded on the M4 Pro and tolerances re-derived from the same 48-draw float32 perturbation study, floored at one float32 unit: no draw moved ll_history[0]. Tested: the canary passes. * Seed the doscaling oracle from a conditioned state From seed 42's normalized init, one sample sits 2.3e-7 from a mixture center; the mu update amplifies the 1e-14 sphere round-off to 1.3e-9 (PyTorch and NumPy differ by 7.9e-10), over the 1e-9 bound. Seed 45 stays 7x under every bound; a new precondition asserts the backends' own disagreement sits 3x below each bound. Tested with the binary. * Re-search the early-merge collapse recipe From its normalized initial A, seed 20's second model keeps gm 1.2e-3 after the second scan instead of collapsing to zero in one iteration. Seed 23 (of 23, 32, 40, 41 in seeds 0-49) collapses as the test needs. Tested with the binary (AMICA_RUN_FORTRAN=1). * Document the normalized initial mixing matrix Changelog entry with the measured behavior change, ADR 0006's remaining difference marked resolved (and ADR 0007's pointer), the differences guide's init wording and the re-searched collapse recipe. mkdocs build --strict passes. * Re-seed the MLX MIR restore test From its normalized initial A, seed 0's discarded step moved the MIR by 2.1e-4 relative (3.4e-5 to 2.1e-3 over 8 perturbed inputs), inside the 1e-4 margin on the CI GPU. Seed 7 moves it by 2.3e-3 to 2.9e-3 on every input. Tested: test_mlx_mir.py passes. * List the re-seeded MLX MIR test in the changelog * Refuse a zero-norm component in normalization normalize_components now raises ValueError on a zero or non-finite row norm instead of returning NaN. The reference path cannot reach it (the unit diagonal keeps every norm >= 1). Tested: zero and NaN rows. * Check a supplied NumPy A at fit start A supplied A (or one left by a previous fit) now raises ValueError on a wrong shape, a non-finite entry or a singular model block, through the shared validate_supplied_mixing. Before, it raised LinAlgError deep in the fit, or (NaN) ended with no stop reason. Tested: all three. * Test NumPy fix_init and the wrapper round trip fix_init (an attribute) starts from the stacked identity and leaves the generator untouched; AMICA save/load restores the PyTorch A bit for bit. Tested: test_initial_mixing.py passes. * Separate the freeze from the collapse figures The 9.0e-4 gm drop is pamica's own fit, whose A is held on those iterations; the 4.2e-3 both sides reach comes from the freeze-free continuation. Docs and test comment now say so. Docs only. * Tighten three comments from the review ADR 0006 marks the init difference as former; the _FLOOR_MARGIN comment names the worse of doscaling on and off; pre_change notes a component is a block row in both layouts. Comments only. * Record the supplied-A check in the changelog * Sphere both oracle sides with the reference S The merged-state oracle now sphers pamica's data with the binary's own sphere, so both sides start from the same bits after the Phase 12 and 13 merge. Tested: gated component-row oracles pass. * Start the pre-344 control from the new init The pre-#344 control class predates #341, so it now starts from the normalized initial A the seeded reference uses. Tested: gated family oracles pass.
* Delete the dead legacy pdf code in numpy_impl/pdf.py choose_pdf_type had no callers, and compute_pdf's pdftype 2-4 branches were unreachable: every caller uses the generalized Gaussian. Tested with the pdf, reference-constant and viz suites. * Correct public docstrings against the epic code Rewrap docstring lines that began with an issue number, which the API pages rendered as headings; state the sphered-space shapes of the wrapper's mixing and unmixing accessors; name restart_error among the degenerate stops; drop the claim that natural gradient alone falls short of the reference; fix the NumPy n_channels and doscaling wording. Checked on the bundled sample (accessor shapes and composition). * Say that maxdecs counts decreases, not consecutive ones The reference resets its decrease count only at each ratchet and when Newton switches on (amica15.f90:1062-1076), as every backend does. * Describe the reference's iteration order in the concept pages How AMICA works now walks one iteration as the code runs it: the E-step, the likelihood-decrease response and stopping checks, the exit before any update, then the update with Newton, the A-freeze windows and the doscaling rescale. The false claim that every iteration raises the likelihood is gone, and the flow figure shows the checks before the update. What is AMICA covers the unit-norm scale convention. * Fix the EEGLAB guide's scalp-map example and multi-model note The variance-order example took scalp maps from the sphered-space get_mixing_matrix; it now uses get_sensor_mixing_matrix. The multi-model note still said the per-model layout differs from the reference's, which #159 and #334 fixed. The guide index lists all four guides. * Correct the backends guide on devices, CUDA float32 and scope Auto device selection picks MPS first and the wrapper redirects it to CPU for float64; the raw AMICATorchNG raises there instead (checked on this host). CUDA float32 is about as fast as float64, not faster, as the same page says above. The native backend is listed, the wrappers' two backends and their parameter-file route are stated, and a stale 'sweep is being finalized' note points to the published sweeps. * Document the raw PyTorch backend's automatic device choice With device=None it picks MPS first and, at the float64 default, raises on an Apple machine; the AMICA wrapper redirects to CPU. Checked on this host. * Update install, accessor and precision notes on the entry pages pamica is on PyPI (0.3.3), so getting started no longer says a release is planned, and says unreleased changes need a source install. The README's get_mixing_matrix comment called the sphered-space matrix scalp maps; the examples now use get_sensor_mixing_matrix for maps. The home page no longer calls float32 faster, the final_ll_ note covers the max_iter case, and the NumPy CLI takes either parameter-file format. * Correct the API overview and backend pages List AMICANative in the public import surface; install extras with uv from PyPI or a source checkout; correct the NumPy page's n_channels claim (the raw PyTorch and MLX constructors take it) and list what that backend lacks; drop the viz page's reference to the closed #159 as an open question; replace em-dashes in list items. * Say the reference compiles Newton off, not on amica15_header.f90 declares do_newton = .false.; only the bundled input.param turns it on. The Newton docstrings and the concept page said the reference runs Newton by default. * Correct the differences guide's Newton row and backend table The reference compiles do_newton off (only the bundled input.param turns it on); row 5 names singular_ll; the backend table gains the variance_order, scoring and wrapper rows and the wrapper .pt save; an obsolete pre-#334 comparison of two retired sharing metrics is replaced by the current single-kernel account; the LLt exception names the last iteration; table cells say none instead of an em-dash. * Bring the validation guide's non-numeric claims up to date The degenerate-fit row still said raw backends were ungated (#306 is done) and named only the likelihood stops; the EEGLAB section still said multi-model files are not in the reference layout. Parity figures are left to #351. * Make the Unreleased changelog one coherent account Group the phase entries by what changed (fitting, MLX and the wrappers, the MNE wrapper, persistence and exports, fixes, removals, docs) and open with a warning that default fits differ from 0.3.3 and why. Drop the entries a later phase superseded (Phase 8's pending #344 note, the polish round's harness and Phase 1's remaining-gap notes), merge the duplicate pdftype-default entries, correct the n_channels claim, and mark measurements taken before later phases. No figure is changed. * Map the shared modules and the epic's fitting changes in AGENTS.md Add rank, schedule, initialization, component_layout and reference_constants (with blocktune, restarts and the params reader) as the shared-decision modules, plus mne_compat, native and metrics; name amica15.f90 as the reference source; add a status line for the reference-fidelity changes and the #306/#339 degenerate-fit extensions; shorten the epic #278 history sentence. * Bring the context plan and progress summary up to date Record epic #324's phases and the reference-fidelity changes, the #306/#339 degenerate-fit extensions and the #334 sharing oracle; close the stale release items (repository transfer, docs deploy, PyPI and Zenodo are done) and the #15 coverage item. * Document the opt-in native-binary tests and MNE's pre-whitener The testing page names AMICA_RUN_FORTRAN=1 and the mne_tests directory; the MNE page's source formula includes the per-channel-type scaling the export applies. * Say shared component, not shared column, in AMICAICA's docstring * Note that doscaling is on by default in What is AMICA * Avoid em-dashes in the plan's updated phase headings * Mention the post-update non-finite stop in How AMICA works * Reword the paper's framing in nuanced terms Drop negation-contrast and candor framing outside the Validation section, count releases as of today (nine, six on PyPI), and describe the MEG status as end-to-end use without measured parity. * Drop candor framing in the differences guide and ADR 0007 * Rebuild paper.pdf * State the complete transform in get_unmixing_matrix The docstring called W @ get_sphere() the full transform; transform also subtracts the mean before sphering and the model center after. The MIR docstrings and the metrics page now call W @ sphere the linear part. Checked on the bundled sample with two models: the full formula matches transform exactly, W @ sphere @ X is off by up to 4.2, and MIR is unchanged by a shift of the data. * Say which default fits #335 and #334 reach, in the changelog The opening warning said the schedule gates and component rows matter only with Newton, rejection, restarts or sharing; they also reach plain default fits in narrow cases (a maxdecs ratchet past newt_start, and the per-component ndtmpsum sum near a grad-norm stop). The stopping iteration's skipped steps are listed in code order (kurtosis switch, share scan, MIR waypoint, rejection), in ADR 0008 too, and the #304 entry says the fit applies the file's settings. * Replace the two em-dashes left in the validation guide Prose only; no figure on those lines changed. * Say that fit applies from_params_file's settings from_params_file stores the file's settings on the instance; fit passes them to the backend constructor as defaults and emits the warning. * State effects directly in the prose this PR added Reword contrast framings in the concept page, home page, viz page and module docstring, and the differences guide's opening, per the style rule: say what the code does and its known effect. --------- Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
* Take the wrapper's fit defaults from the backend AMICA.fit restated max_iter/lrate/do_mean/do_sphere/do_newton, and its lrate (0.05, runamica15.m's) differed from every backend's 0.1. It now reads them from the selected backend's signatures, so AMICA and AMICAICA fits without lrate run at 0.1. Tests for torch and MLX: resolved defaults, a default wrapper fit equals the default backend fit, and one set of constructor defaults across the backends. * Move the MPS/float64 device fallback into AMICATorchNG AMICATorchNG(n_channels) with default arguments raised ValueError on Apple Silicon: device=None picked MPS, which has no float64. The constructor now uses the CPU for that case and logs a warning; the wrapper and AMICA.load pass device through and drop their copy. An explicit device='mps' at float64 still raises. Torch-only (MLX has no device choice). Tested on MPS: raw construction and fit, float32 auto, explicit MPS, load with device=None. * Remove the dead pamica/data package-data entry No pamica/data directory exists. Built with uv build: the wheel's file list is unchanged, it imports from a clean venv, and numpy_impl/params.json is inside. * Document pamica's defaults and the backend device choice Add a defaults table to the differences guide: pamica against the compiled amica15 binary and EEGLAB's runamica15.m, why pamica follows the binary, and how to reproduce an EEGLAB run. Update the pages that stated the wrapper's lrate or the raw backend's MPS raise. mkdocs build --strict passes. * Add the Phase 17 changelog entries The wrapper lrate change joins the 'Default fits differ from 0.3.3' warning; entries for the backend device fallback and the package-data cleanup. * Give AMICA_NumPy the shared default settings params.json set do_newton on and max_iter 2000, the only shared settings where NumPy differed from AMICATorchNG; it now sets them off and 100, as does the constructor's max_iter fallback (reached by the CLI). Fits shorter than newt_start are byte-identical. New test_default_settings.py holds NumPy and MLX to every shared AMICATorchNG default (the MLX check moves there from test_wrapper_backends.py). * Correct the degenerate-fit test's measured stop The comment said torch stops on nan_ll at iteration 3; both backends stop on nan_params with iteration 2 in all 16 measured runs, at either lrate. The test asserts only a degenerate stop. * Document the NumPy backend's shared defaults Defaults section, validation guide's stops table, NumPy API page, and a changelog entry (NumPy joins the opening warning's defaults line). * Give AMICANative pamica's shared defaults The engine wrote the bundled input.param's values (lrate 0.05, Newton on, max_iter 2000, block_size 512, ...). It now takes every setting the binary has a keyword for from AMICATorchNG's signature through FORTRAN_TO_PAMICA_KEY; binary-only knobs are unchanged. The default block is pamica's 8192 samples split over max_threads and capped at the data, since the binary's block_size is per thread and a block longer than the data gives an all-NaN fit (#292). The native oracles now start from the bundled input.param itself (byte-identical files). Tested with the v0.3.3 binary: native, oracle and default-settings tests pass. * Document AMICANative's shared defaults AMICANative joins the defaults table's pamica column; the section and the native API page cover the settings it cannot write, the per-thread block size, and running the binary as an input.param configures it. The changelog records the behavior change. * Say where AMICANative's default pcakeep comes from
* Add issue-351 Newton independent-seed study script * Add hallu driver for the Newton seed study * Record harness, same-start and fixture parity runs * Record share_comps counts and float32 agreement runs * Update share_comps merge counts in differences guide * Re-state the harness rows with same-start parity * Add an MLX float32 option to the Newton seed study * Mark the issue-27 and issue-51 records as pre-epic * Update the harness figures in AGENTS and progress summary * Add ensemble scripts and dimsweep LL records * Add the hallu CUDA dimsweep LL record * Re-state the cross-backend LL table with the chaos caveat * Reword the harness and cross-backend notes * Pin the ensemble lrate in reproduce_table1 * Add basin and pre-epic ensemble scripts * Record the bundled tier, basin, ensemble and MLX Newton runs * Add the equivalence-margin check on saved ensembles * Record the seeded keep_best ensemble * Add the re-measured keep_best figures to ADR 0003 * Name the pre-change code precisely in validation guide * Regenerate the multi-model ensemble figure * Rebuild paper.pdf * Re-state the bundled single-model and multi-model results * Update keep_best and multi-model figures in AGENTS and guides * Update the fixture parity figures in AGENTS and progress summary * Record the bundled tier wall-clock on the M4 Pro * Update the cross-backend LL claim in the backends guide * Record the Amari pairing-correction measurement * Update paper Table 1 bundled rows and the float32 paragraph * Clarify the benchmark excerpt in the paper * Rebuild paper.pdf * Record the gated oracle and full-suite runs * Add the Phase 15 re-measurement to the changelog * Record the Newton independent-seed fits * Record the same-session 70-channel timing check * Record the unseeded-start check of the issue-145 binary * Mark the sections of the validation guide that predate the epic * Pin every reference setting in the Table 1 and run scripts * Re-state the Newton-enabled basin results * Give both matching seeds in the paper * Rebuild paper.pdf * Document the ensemble's invsigmin difference * Smooth the ensemble protocol sentence * Address the docs-claims review findings * Address the paper review findings * Give each ensemble distribution its own line style * Label the external single-model figures as pre-epic * Reflow the README parity sentence * Record the external tier run * Fill the external-tier figures and fix rounding slips * Rebuild paper.pdf * Save the pinned parameter files before Phase 17 * Re-check the reference pins after Phase 17 * Pin do_approx_sphere in the reference settings * Record the test runs after the Phase 17 merge * Add the issue-351 findings record * Reword two remaining contrast phrasings in the paper * Rebuild paper.pdf --------- Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Filter CI on code paths (docs, paper, paper.pdf, metadata and non-Python .context records skip; .context Python files and the tested ensemble.npz still run), and commit the rebuilt paper.pdf back only on dev and main. Checked with actionlint.
Epic #324: first-class MLX and reference-faithful fitting
auto-bump-dev.yml can push to dev while the PDF builds, which rejected the plain push after the epic #324 merge. Checked with actionlint.
* Date the changelog's release headings Add each release's tag date to its heading, describe the Unreleased convention in the intro, and record the changelog mechanism. * Point a root CHANGELOG.md and PyPI at the changelog * Extract a release's changelog section as release notes scripts/changelog_section.py prints one release's section with its relative docs links made absolute, and --check gates a release. Tested on the real changelog (11 tests). * Publish release notes from the changelog and check entries auto-tag.yml uses the release's changelog section as the notes, with generated PR notes appended. changelog.yml requires an entry for package-code PRs into dev (skip-changelog label to opt out) and the release section for PRs into main. Checked with actionlint. * Record the changelog practice in the rules * Date releases by their UTC publication date v0.1.3, v0.2.0 and v0.3.2 carried their tagged commit's date; use the GitHub release date (or the tag date where no release exists), and state the convention in the rule. * Leave code untouched and rewrite reference links in release notes The extractor skips code spans and fenced blocks, ignores headings inside fences, and rewrites reference-style link definitions. Seven edge-case tests (18 total) cover them.
Rename the changelog's Unreleased section to 0.4.0 (2026-09-24) with a parity summary, and set the version to 0.4.0.dev0 with sync_version.py so the release chain tags 0.4.0.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Release 0.4.0. Merge this with a regular merge commit.
auto-tag.ymlthen strips.devN, tagsv0.4.0and creates the GitHub release, using the changelog's 0.4.0 section as the notes. That startspublish.yml(PyPI),release-binaries.ymlandsync-dev.yml.0.4.0 at a glance
MLX becomes a first-class backend, reachable through every wrapper feature (epic #324, completing the raw backend of epic #278),
and every backend's fitting follows the Fortran reference more closely.
Against the pinned v0.3.3 reference binary, single-model fits of a 70-channel EEG recording (5 seeds, 2000 iterations)
now agree with it to 6e-6 in log-likelihood, with a mean matched component correlation of 0.9996
(Validation & Parity).
The full release note is the 0.4.0 section of the changelog. It covers:
Pre-release checks
0.4.0.dev0(set in Prepare the 0.4.0 release #367).CITATION.cffand.zenodo.jsonagree.## 0.4.0 - 2026-09-24and no## Unreleased. This PR's Changelog check verifies both on the merged tree.deletion,non_fast_forward,pull_request; dev has none. TheAUTO_TAG_PATsecret is present.