This project is experimental research and a work in progress.
The goal is to bring local, error-driven learning to full-size transformers, then connect that learning to a symbolic knowledge store. This repository explores that direction with GPT-2 at 124 million parameters: predictive-coding training in PyTorch, continual-learning experiments, and inference through the included MORK Rust engine.
Predictive coding gives each transformer block an error state. Those states settle against an objective, and each block updates its weights from its local settled target. Homotopy distillation gradually increases the settling depth while preserving the pretrained model's predictions. It provides a working starting point for studying how transformers can adapt through local learning rules.
The longer-term aim is a transformer that learns alongside symbolic knowledge inside MORK. The current experiments develop the learning rule and the inference adapter needed to pursue that goal.
Reproduce the report · Method · Research report · All results
The main run completed 50 million OpenWebText training tokens. Across its 51 milestones, the largest student–teacher prompt KL was 3.1 × 10⁻⁵ nats. The plots show prompt agreement and held-out perplexity throughout the schedule, using the fixed 32-token evaluation windows.
A separate terminal evaluation used context length 512 and 409,600 evaluated tokens per dataset. On unseen OpenWebText, student perplexity matched the teacher's to 0.0032%. Lower perplexity is better.
| Evaluation data | Teacher perplexity | Student perplexity | Change |
|---|---|---|---|
| Unseen OpenWebText | 24.3079 | 24.3072 | −0.0032% |
| Code | 23.7360 | 23.9915 | +1.08% |
| Legal text | 20.4925 | 20.4821 | −0.051% |
| Biomedical abstracts | 23.7535 | 23.7521 | −0.0059% |
Along the main schedule, the local block-gradient norm grew from 0.10 to 0.98 times the backpropagation gradient norm, with minimum block cosine similarity above 0.9986. The learning signal changes in magnitude while its direction remains aligned. Training measurements · Terminal evaluation
At context length 512, GPT-2 has 4.72 million error cells per sample. In the recorded terminal settle, the largest 5% of cells carry about 90% of squared error mass. A cross-entropy settle also shows this concentration, with 91.6% in the largest 5%. That structure motivates work on making settling cheaper by concentrating computation where the errors are largest.
Frontier measurements and replay commands
The continual-learning experiments train on code, then legal text, then biology, with a 4.8-million-token budget per domain. An evidence gate adjusts local weight updates using accumulated error statistics.
Across the three recorded seeds, gated predictive coding achieved lower relative forgetting and lower final mean perplexity than both ungated predictive coding and backpropagation on this curriculum. Each line connects the same seed across methods.
Per-domain results, settings, and controls
The inference adapter represents transformer weights and activations as tensor expressions in MORK and executes the forward operations through its Rust tensor kernels. It includes full-forward and incremental-decode comparisons against NumPy references. The recorded distilled-model check uses all 12 layers and matches both references for all three generated tokens.
The next research steps are to reduce settling cost, extend the comparisons of local learning rules, and connect weight updates to the MORK inference substrate.
| Task | Start here |
|---|---|
| Reproduce a reported experiment | Reproduction guide |
| Train or resume a model | Training guide |
| Run MORK inference | Inference guide |
| Compare experiments | Results |
| Use saved weights | Model files |
| Understand the method | Research and report |
| Prepare cluster data | Cluster guide |
From the repository root:
python3.11 -m venv .venv
.venv/bin/python -m pip install -r cluster/requirements.txt \
--extra-index-url https://download.pytorch.org/whl/cu128
.venv/bin/python -m pip install --no-deps -e .
.venv/bin/hdpc-distill --run-name smoke --steps 200 --seq-len 16 --batch-size 1Training uses PyTorch. Runs and checkpoints are written to runs/.
For real training data, see token shards and data modes.
Install Rustup and the native build prerequisites, then:
bash scripts/build_mork.sh
.venv/bin/hdpc-mork-decode --prompt Hello --n-gen 3 --stripsMORK performs inference; predictive-coding training runs in PyTorch.
The command uses pretrained GPT-2 by default. Pass --weights to use
distilled weights.
src/hdpc/ Training, checkpoints, and diagnostics
src/hdpc/mork/ MORK inference adapter and NumPy reference
vendor/ MORK and PathMap Rust source
tests/ Training and inference tests
results/ Measurements grouped by experiment
reports/ Research report and figures
docs/ Usage guides and research notes
cluster/ Data preparation and offline setup
scripts/ Engine build and report generation
Weights, datasets, environments, and generated runs are excluded from Git.
OMP_NUM_THREADS=16 MKL_NUM_THREADS=16 .venv/bin/python -m pytest testsBuild MORK before running the inference tests. See testing for focused checks and the exact fp32 environment.


