Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 13 additions & 7 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,23 +3,29 @@
## Project Context
**Purpose:** Python implementation of AMICA (Adaptive Mixture Independent Component Analysis) that reproduces the results of the reference Fortran binary. Targets EEG/EMG source separation with GPU/MPS/CPU support.
**Tech Stack:** Python 3.12+, PyTorch (primary backend, MPS/CUDA/CPU), NumPy/SciPy (legacy backend), matplotlib. Reference implementation is Fortran (`amica17.f90`, `funmod2.f90`).
**Architecture:** The scikit-learn-style `AMICA` interface wraps `AMICATorchNG` (`torch_impl/amica_torch_ng.py`), the natural-gradient EM port that reaches Fortran parity (Newton, exact-EM mixture updates, symmetric-ZCA sphere, Jacobian LL). This is the single PyTorch backend: the earlier Adam/autograd backends (`AMICATorch`, `AMICATorchV2`) and their mixture/optimizer/PDF helper modules were removed in issue #32 as superseded. The legacy NumPy implementation (`pyAMICA.py`, retained as `AMICA_NumPy`) carries the same parity fixes plus baralpha and outlier rejection. Correctness is defined by parity with the Fortran binary, validated by `validate_implementations.py`.
**Architecture:** The scikit-learn-style `AMICA` interface wraps `AMICATorchNG` (`torch_impl/core.py`), the natural-gradient EM port that reaches Fortran parity (Newton, exact-EM mixture updates, symmetric-ZCA sphere, Jacobian LL). This is the single PyTorch backend: the earlier Adam/autograd backends (`AMICATorch`, `AMICATorchV2`) and their mixture/optimizer/PDF helper modules were removed in issue #32 as superseded. The legacy NumPy implementation (`numpy_impl/core.py`, retained as `AMICA_NumPy`) carries the same parity fixes plus baralpha and outlier rejection. Correctness is defined by parity with the Fortran binary, validated by `validate_implementations.py`.

## Architecture Map
```
pyAMICA/
├── amica.py # Main scikit-learn-style AMICA interface (wraps AMICATorchNG)
├── __init__.py # Exposes AMICA (PyTorch) and AMICA_NumPy (legacy)
├── __init__.py # Exposes AMICA (PyTorch), AMICA_NumPy (legacy), numpy_impl, torch_impl
├── torch_impl/ # PyTorch backend
│ ├── amica_torch_ng.py # Natural-gradient EM port (AMICATorchNG); Fortran-parity, sole backend
│ ├── core.py # Natural-gradient EM port (AMICATorchNG); Fortran-parity, sole backend
│ └── utils.py # Preprocessing (sphering, PCA), device selection
├── pyAMICA.py, amica_*.py # Legacy NumPy implementation + CLI/data/viz/pdf/newton
├── numpy_impl/ # Legacy NumPy reference (topic-named modules, issue #34)
│ ├── core.py # AMICA_NumPy; newton.py, pdf.py, data.py, load.py, viz.py, utils.py, cli.py
│ └── ...
├── amica17.f90, funmod2.f90 # Fortran reference source (read-only, for parity)
├── sample_data/ # Sample EEG data + Fortran binary (amica15mac)
└── tests/ # Tests, incl. tests/torch_tests/ (vs-Fortran parity)

validate_implementations.py # Runs both implementations, Hungarian component matching, reports
```
Module names are topic-based (`core`/`newton`/`pdf`/`data`/... under `numpy_impl/`,
`core`/`utils` under `torch_impl/`); the old `pyAMICA.py`/`amica_*.py`/`amica_torch_ng.py`
prefixes were dropped in issue #34. The public import surface is stable:
`from pyAMICA import AMICA, AMICA_NumPy, AMICATorchNG`.

## Environment Setup
Canonical environment is **UV** (per global standards). The PyTorch stack is declared in
Expand All @@ -36,7 +42,7 @@ computes in float64 for Fortran parity, which MPS cannot represent, so parity ru

## Key Files
- **Main interface:** `pyAMICA/amica.py` (thin wrapper over `AMICATorchNG`)
- **PyTorch backend:** `pyAMICA/torch_impl/amica_torch_ng.py` (`AMICATorchNG`, natural-gradient EM,
- **PyTorch backend:** `pyAMICA/torch_impl/core.py` (`AMICATorchNG`, natural-gradient EM,
Fortran-parity; ADR `.context/decisions/0001-torch-backend-natural-gradient-em.md`). This is the
only PyTorch backend; the basic `AMICATorch`/`AMICATorchV2` paths were removed in #32.
- **Validation harness:** `validate_implementations.py`
Expand All @@ -49,15 +55,15 @@ computes in float64 for Fortran parity, which MPS cannot represent, so parity ru
and positive-definite (issue #24).
- Validation harness runs both implementations (NG + NumPy) and matches components via the Hungarian
algorithm.
- Newton and exact-EM updates are implemented in `AMICATorchNG` and the legacy NumPy `pyAMICA.py`
- Newton and exact-EM updates are implemented in `AMICATorchNG` and the legacy NumPy `numpy_impl/core.py`
(both Fortran-faithful). Adaptive PDF (#26) is DONE (all five `pdftype` families + ext-Infomax
switcher); full multi-model matching (#27) is validated by distributional equivalence.
- See `FEATURE_PARITY.md`, `MIGRATION_PLAN.md`, and `PROGRESS_SUMMARY.md` for detailed roadmaps.

## Known Issues (parity blockers)
**Single-model parity: DONE (#24).** The natural-gradient A-update transpose fix (plus exact-EM
mixture updates, digamma rho update, symmetric-ZCA sphere, Jacobian LL) brought both `AMICATorchNG`
and the legacy NumPy `pyAMICA.py` to Fortran's solution (LL ~ -3.40, Hungarian-matched component
and the legacy NumPy `numpy_impl/core.py` to Fortran's solution (LL ~ -3.40, Hungarian-matched component
correlation ~0.997, > 0.95 gate cleared; root cause in `.context/issue-24/`). Also resolved: Newton
stability (posdef, 0 fallbacks), backend consolidation (#32/#31), NumPy CLI save/load format (#30),
NG save/load persistence (#36), and the degenerate-fit contract (#50: the `AMICA` wrapper marks a
Expand Down
4 changes: 2 additions & 2 deletions MIGRATION_PLAN.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,7 +53,7 @@ Replace the NumPy-based AMICA implementation with the PyTorch version to leverag
```

2. **Update CLI**
- Modify `amica_cli.py` to use PyTorch implementation
- Modify `numpy_impl/cli.py` to use PyTorch implementation
- Keep same interface for backward compatibility

3. **Update Tests**
Expand Down Expand Up @@ -127,7 +127,7 @@ model = AMICA(n_channels=32, device='cpu')
pytest pyAMICA/tests/

# Save reference results
python -m pyAMICA.amica_cli sample_params.json --outdir numpy_reference
python -m pyAMICA.numpy_impl.cli sample_params.json --outdir numpy_reference
```

### After Migration
Expand Down
2 changes: 1 addition & 1 deletion PROGRESS_SUMMARY.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@
- Documented critical missing features and priorities

### 5. Natural-gradient EM backend at Fortran parity ✅ (epic #9 / issue #24)
- Added `AMICATorchNG` (`torch_impl/amica_torch_ng.py`); `AMICA` wraps it directly
- Added `AMICATorchNG` (`torch_impl/core.py`); `AMICA` wraps it directly
- Root-caused the parity gap: the natural-gradient A-update was transposed / multiplied on the
wrong side (proven machine-exact); plus exact-EM mixture updates, the digamma rho update, the
symmetric-ZCA sphere, the output transpose, and the NumPy Jacobian LL
Expand Down
16 changes: 9 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,7 +43,7 @@ pip install -e .

2. Run AMICA:
```bash
python amica_cli.py params.json --outdir results
python -m pyAMICA.numpy_impl.cli params.json --outdir results
```

## Parameters
Expand Down Expand Up @@ -108,12 +108,14 @@ python amica_cli.py params.json --outdir results

## File Structure

- `amica.py`: Main AMICA implementation
- `amica_pdf.py`: PDF type implementations
- `amica_newton.py`: Newton optimization
- `amica_data.py`: Data loading/preprocessing
- `amica_cli.py`: Command-line interface
- `params.json`: Example parameter file
- `amica.py`: Main scikit-learn-style AMICA interface (PyTorch backend)
- `torch_impl/core.py`: PyTorch natural-gradient EM backend (`AMICATorchNG`)
- `numpy_impl/core.py`: Legacy NumPy reference (`AMICA_NumPy`)
- `numpy_impl/pdf.py`: PDF type implementations
- `numpy_impl/newton.py`: Newton optimization
- `numpy_impl/data.py`: Data loading/preprocessing
- `numpy_impl/cli.py`: Command-line interface
- `numpy_impl/params.json`: Default/example parameter file

## Output Files

Expand Down
19 changes: 7 additions & 12 deletions pyAMICA/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -16,23 +16,18 @@
from .version import __version__
from .amica import AMICA
from .torch_impl import AMICATorchNG
from . import amica_utils
from . import amica_data
from . import amica_newton
from . import amica_pdf
from . import amica_viz
from . import numpy_impl
from . import torch_impl

# Legacy NumPy implementation (deprecated)
from .pyAMICA import AMICA as AMICA_NumPy
# Legacy NumPy reference implementation (topic-named modules under numpy_impl/,
# issue #34); AMICA_NumPy is its scikit-learn-style interface.
from .numpy_impl import AMICA as AMICA_NumPy

__all__ = [
"AMICA",
"AMICATorchNG",
"AMICA_NumPy",
"amica_utils",
"amica_data",
"amica_newton",
"amica_pdf",
"amica_viz",
"numpy_impl",
"torch_impl",
"__version__",
]
11 changes: 11 additions & 0 deletions pyAMICA/numpy_impl/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,11 @@
"""Legacy NumPy reference implementation of AMICA.

Topic-named modules (``core``, ``newton``, ``pdf``, ``data``, ``load``, ``viz``,
``utils``, ``cli``), renamed from the old ``pyAMICA.py``/``amica_*.py`` sprawl in
issue #34. The scikit-learn-style :class:`~pyAMICA.numpy_impl.core.AMICA` is also
exposed at the package root as ``pyAMICA.AMICA_NumPy``.
"""

from .core import AMICA

__all__ = ["AMICA"]
6 changes: 3 additions & 3 deletions pyAMICA/amica_cli.py → pyAMICA/numpy_impl/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,7 +10,7 @@
4. Result saving

Example usage:
python -m pyAMICA.amica_cli params.json --outdir results --seed 42 # Use -m flag to run as module
python -m pyAMICA.numpy_impl.cli params.json --outdir results --seed 42 # Use -m flag to run as module

The parameter file should be in JSON format and must include:
- files: List of binary data files to process
Expand All @@ -31,8 +31,8 @@
from pathlib import Path
from typing import Dict, Any

from .pyAMICA import AMICA
from .amica_data import load_multiple_files
from .core import AMICA
from .data import load_multiple_files


def parse_args() -> argparse.Namespace:
Expand Down
14 changes: 7 additions & 7 deletions pyAMICA/pyAMICA.py → pyAMICA/numpy_impl/core.py
Original file line number Diff line number Diff line change
Expand Up @@ -57,10 +57,10 @@

See Also
--------
amica_pdf : PDF implementations
amica_utils : Utility functions
amica_viz : Visualization tools
amica_cli : Command-line interface
pdf : PDF implementations
utils : Utility functions
viz : Visualization tools
cli : Command-line interface

References
----------
Expand All @@ -79,7 +79,7 @@
from pathlib import Path
from typing import Dict, Optional, Tuple
from tqdm import tqdm
from .amica_utils import (
from .utils import (
gammaln,
determine_block_size,
identify_shared_components,
Expand Down Expand Up @@ -385,7 +385,7 @@ def fit(self, data: Optional[np.ndarray] = None) -> "AMICA":
"(a length mismatch would silently truncate to the "
"shorter list via zip())."
)
from .amica_data import load_multiple_files
from .data import load_multiple_files

data = load_multiple_files(
self._config_files, self._config_data_dim, self._config_field_dim
Expand Down Expand Up @@ -1354,7 +1354,7 @@ def _write_results(self):
"""Write current results to disk in the Fortran AMICA binary format.

Writes raw little-endian float64 (and int32 ``comp_list``) files with no
extension, in the layout that ``amica_load.loadmodout`` reads (and that
extension, in the layout that ``load.loadmodout`` reads (and that
``load_results`` reads back), so pyAMICA output is loadable by the same
reader as the Fortran reference (issue #30).

Expand Down
4 changes: 2 additions & 2 deletions pyAMICA/amica_data.py → pyAMICA/numpy_impl/data.py
Original file line number Diff line number Diff line change
Expand Up @@ -201,9 +201,9 @@ def load_results(indir: str, compressed: bool = False) -> dict:
Load saved AMICA results from disk.

Reads the raw little-endian binary files written by ``AMICA._write_results``
-- the Fortran AMICA output layout that ``amica_load.loadmodout`` and the
-- the Fortran AMICA output layout that ``load.loadmodout`` and the
reference binary use (issue #30) -- and returns them in AMICA's internal
array shapes, which the ``amica_viz`` helpers consume.
array shapes, which the ``viz`` helpers consume.

Parameters
----------
Expand Down
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
File renamed without changes.
4 changes: 2 additions & 2 deletions pyAMICA/amica_viz.py → pyAMICA/numpy_impl/viz.py
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@
from pathlib import Path
from typing import Optional, Union, Tuple

from .amica_data import load_results
from .data import load_results


def plot_convergence(
Expand Down Expand Up @@ -264,7 +264,7 @@ def plot_pdf_fits(
Legacy no-op kept for signature compatibility; results are read from the
raw Fortran binary format, never compressed .npz (issue #30).
"""
from .amica_pdf import compute_pdf
from .pdf import compute_pdf

results = load_results(results_dir, compressed)

Expand Down
6 changes: 3 additions & 3 deletions pyAMICA/tests/test_amica.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,9 @@
from pathlib import Path
import unittest

from pyAMICA.amica_data import load_data_file, preprocess_data
from pyAMICA.amica_pdf import compute_pdf
from pyAMICA.amica_newton import compute_newton_direction
from pyAMICA.numpy_impl.data import load_data_file, preprocess_data
from pyAMICA.numpy_impl.pdf import compute_pdf
from pyAMICA.numpy_impl.newton import compute_newton_direction


class TestAMICA(unittest.TestCase):
Expand Down
4 changes: 2 additions & 2 deletions pyAMICA/tests/test_pyAMICA.py
Original file line number Diff line number Diff line change
Expand Up @@ -8,8 +8,8 @@

import pyAMICA
from pyAMICA import AMICA_NumPy as AMICA
from pyAMICA.amica_data import load_data_file, preprocess_data
from pyAMICA.amica_pdf import compute_pdf
from pyAMICA.numpy_impl.data import load_data_file, preprocess_data
from pyAMICA.numpy_impl.pdf import compute_pdf

# Setup test data path
data_path = op.join(pyAMICA.__path__[0], "data")
Expand Down
26 changes: 13 additions & 13 deletions pyAMICA/tests/test_sample_data.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,8 @@

import pyAMICA
from pyAMICA import AMICA_NumPy as AMICA
from pyAMICA.amica_data import load_data_file
from pyAMICA.amica_load import loadmodout
from pyAMICA.numpy_impl.data import load_data_file
from pyAMICA.numpy_impl.load import loadmodout

# Setup sample data paths
sample_data_path = op.join(pyAMICA.__path__[0], "sample_data")
Expand Down Expand Up @@ -75,7 +75,7 @@ def test_sample_data_scikit(tmp_path):
def test_sample_data_cli():
"""Full CLI-vs-Fortran integration test (issue #30 format + #39/#41 stability).

Runs the real amica_cli entrypoint for the full 2000-iter sample config and
Runs the real cli entrypoint for the full 2000-iter sample config and
Hungarian-matches the loadmodout-read W against the Fortran reference with a
strict all-32-components > 0.8 gate. A fixed seed is pinned so the test does
not depend on a random init (see #39). Currently xfail on a residual
Expand All @@ -85,7 +85,7 @@ def test_sample_data_cli():
import subprocess
import sys

# amica_cli.py uses relative imports, so it must be run as a module
# cli.py uses relative imports, so it must be run as a module
# (see its module docstring), not as a direct script path.
test_outdir = Path("test_output")

Expand All @@ -94,7 +94,7 @@ def test_sample_data_cli():
[
sys.executable,
"-m",
"pyAMICA.amica_cli",
"pyAMICA.numpy_impl.cli",
sample_params_file,
"--outdir",
str(test_outdir),
Expand Down Expand Up @@ -243,8 +243,8 @@ def test_cli_output_format_roundtrip(tmp_path):
import matplotlib

matplotlib.use("Agg")
from pyAMICA.amica_data import load_results
from pyAMICA import amica_viz
from pyAMICA.numpy_impl.data import load_results
from pyAMICA.numpy_impl import viz

data = load_data_file(eeglab_data_file, 32, 30504, dtype=np.float32).astype(
np.float64
Expand Down Expand Up @@ -286,9 +286,9 @@ def test_cli_output_format_roundtrip(tmp_path):
np.testing.assert_array_equal(r["comp_list"], model.comp_list)

# viz helpers run without error on the loaded results.
amica_viz.plot_convergence(str(outdir))
amica_viz.plot_components(str(outdir), data=None, max_comps=3)
amica_viz.plot_pdf_fits(str(outdir), data, max_comps=2)
viz.plot_convergence(str(outdir))
viz.plot_components(str(outdir), data=None, max_comps=3)
viz.plot_pdf_fits(str(outdir), data, max_comps=2)


@pytest.mark.skipif(not op.exists(eeglab_data_file), reason="sample data missing")
Expand Down Expand Up @@ -413,7 +413,7 @@ def test_check_convergence_ratchets_lrate_on_decrease():

@pytest.mark.skipif(not op.exists(eeglab_data_file), reason="sample data missing")
def test_cli_subprocess_output_loadable(tmp_path):
"""The actual amica_cli entrypoint writes loadmodout-readable output.
"""The actual cli entrypoint writes loadmodout-readable output.

Runs the CLI as a module on a short, stable config (real sample data) and
confirms loadmodout reads the result -- the direct regression for the #30
Expand Down Expand Up @@ -441,12 +441,12 @@ def test_cli_subprocess_output_loadable(tmp_path):
params_file.write_text(json.dumps(params))
outdir = tmp_path / "cli_out"

# amica_cli uses relative imports, so run it as a module from the repo root.
# cli uses relative imports, so run it as a module from the repo root.
subprocess.run(
[
sys.executable,
"-m",
"pyAMICA.amica_cli",
"pyAMICA.numpy_impl.cli",
str(params_file),
"--outdir",
str(outdir),
Expand Down
4 changes: 2 additions & 2 deletions pyAMICA/tests/torch_tests/test_ng_backend.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,8 +14,8 @@
import torch

from pyAMICA.torch_impl import AMICATorchNG
from pyAMICA.torch_impl.amica_torch_ng import _KEEP_BEST_TOL
from pyAMICA.pyAMICA import AMICA as AMICA_NumPy
from pyAMICA.torch_impl.core import _KEEP_BEST_TOL
from pyAMICA.numpy_impl.core import AMICA as AMICA_NumPy

SAMPLE_DIR = Path(__file__).resolve().parents[2] / "sample_data"
DATA_FILE = SAMPLE_DIR / "eeglab_data.fdt"
Expand Down
2 changes: 1 addition & 1 deletion pyAMICA/tests/torch_tests/test_ng_pdf_families.py
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@
import torch

from pyAMICA.torch_impl import AMICATorchNG
from pyAMICA.torch_impl.amica_torch_ng import _log_pdf_and_deriv, _score
from pyAMICA.torch_impl.core import _log_pdf_and_deriv, _score

SAMPLE_DIR = Path(__file__).resolve().parents[2] / "sample_data"
DATA_FILE = SAMPLE_DIR / "eeglab_data.fdt"
Expand Down
2 changes: 1 addition & 1 deletion pyAMICA/torch_impl/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
with support for CUDA, ROCm, and Apple Silicon (MPS) backends.
"""

from .amica_torch_ng import AMICATorchNG
from .core import AMICATorchNG
from .utils import setup_device, check_numerical_stability

__all__ = [
Expand Down
Loading
Loading