A research project building models that compress long context into latent concept vectors, reason in that concept space, and decode back to text or other modalities.
Author: Krzysztof Sopyla β ai.ksopyla.com Β· GitHub Β· LinkedIn Project page: ai.ksopyla.com/projects/concept-encoder Experiments: wandb.ai/ksopyla/MrCogito
Today's language models reason by expanding thought into text tokens. That works, but it is an extremely narrow channel: every intermediate idea has to be serialized into a vocabulary item, attended over again, and paid for again in the context window.
This project explores a different primitive:
Compress input into dense latent concepts, refine those concepts recursively, and generate output from the refined concept state.
The end goal is a foundation-model architecture defined by four bold targets:
- π§ Reasoning in latent space. The model should spend most of its compute refining continuous concept states, not just emitting visible chain-of-thought tokens.
- π 10M-token context as the north star. Reached not by forcing full self-attention to scale forever, but by compressing long streams into a far smaller set of concept vectors.
- π One concept substrate across modalities. Text, audio, and eventually other inputs map into the same reasoning space, with modality-specific adapters and decoders only at the edges.
- π€ A path to latent agent communication. Once concept vectors are a reliable semantic interface, models can exchange concepts directly instead of talking through text. That multi-agent layer lives outside this repository β but it is the foundation this work is meant to enable.
The core bet is simple: if concept vectors can carry rich semantic and generative state, long-context reasoning becomes cheaper, more inspectable, and more composable than token-only reasoning.
The frontier's answer to long context has been more of the same: better hardware, clever attention approximations, bigger clusters. I want a different bet β not making self-attention cheaper, but asking whether the model needs it at all for most of its reasoning.
A few older ideas I think were ahead of their time, and that I believe belong together:
- Perceiver1 β a small set of latent queries can cross-attend to a huge input and decouple compute from sequence length. Better parameter and compute utilization through cross-attention β but used for perception, not as the seat of reasoning.
- Embedding space is wildly underused. A single dense vector can be made to hold ~1,500 tokens of recoverable text2, yet we still spend enormous capacity on giant vocabularies and wide token embeddings. I believe most of that capacity goes to surface form, not meaning β the latent space is the asset we're not exploiting.
- Reasoning can move into latent space. Looping the same weights3 and predicting continuous vectors instead of discrete tokens4 let a model "think" beyond the token channel and punch above its parameter count5.
My bet is that combining these into one end-to-end trained architecture β a cross-attention concept bottleneck plus a recursive reasoning core operating in latent space β opens a different paradigm than the scaling race. Cross-attention bottlenecks themselves aren't new (Flamingo6, Large Concept Models7), but they're treated as a means to an end, not as the primary locus of reasoning. The question I keep asking is sharper:
Can a small set of concept vectors become the model's working memory β the thing it actually reasons over?
If yes, the payoff compounds: O(CΒ·N) long context with C βͺ N, test-time compute scaling by running more refinement steps without retraining, modality-agnostic reasoning over a shared concept space, and an internal state small enough to actually inspect. It also opens a channel the token interface can't: letting agents talk to each other in concept space instead of text β early work shows passing latent state between models is both faster and more accurate than exchanging tokens89. It may not work β it's a genuine research bet β but I think it's the more interesting direction, and the one the scaling race is leaving unexplored.
The architecture follows an encode β reason β decode pattern, with the concept space at its center:
Input tokens / frames (N)
β
βΌ
Encoder cross-attention compresses N inputs into C concepts (C βͺ N)
β
βΌ
Reasoning core optional recursive refinement over the concepts
β
βΌ
Decoder generates text, speech, or another modality from concepts
A small set of learned concept queries cross-attends to the full input sequence. This shifts the dominant attention pattern from O(NΒ²) token-to-token attention to O(CΒ·N) concept-to-token attention β and the concept count C grows far more slowly than the input length N.
Input length N |
Concepts C |
Attention savings vs. full self-attention |
|---|---|---|
| 512 | 128 | ~4Γ |
| 32K | 2K | ~16Γ |
| 1M | 8K | ~128Γ |
| 10M | hierarchical | north-star target β needs hierarchical concepts, recurrence, and memory |
This is not just an efficiency trick. The bottleneck is meant to force abstraction: concepts become the model's working memory, the object reasoning operates on, and the interface decoders read from.
The project explores the same concept-bottleneck idea in two complementary settings:
- From-scratch concept encoder (
ConceptEncoder/ PerceiverβBiXT family) β train the encoder, reasoning surface, and decoder end-to-end. Useful for objective and geometry studies (prefixβsuffix AR, windowed decode, STS-B probes). - Pretrained LM + concept workspace (
BackboneConceptLM) β freeze a strong token model (Gemma-3-1B), adapt it lightly with LoRA, and attach a small set of shared depth-recurrent concepts that read from and write into selected backbone depths as the sequence is consumed in blocks (Ktokens β one concept update). The backbone keeps local fluency; the concepts are meant to carry multi-block state.
The Gemma path asks a sharper mechanism question: when next-token loss can already be solved from local context, do concepts ever become causally necessary? Recent long-context training (seq 4096, long-document mix, ~1B tokens) cleared that gate β concept ablations hurt beyond-local positions by a large margin while geometry stayed healthy. That does not close the from-scratch path or latent-reasoning / diffusion threads; it marks one validated regime worth exploring further. Details live in the agenda and E16b run report.
This is an active research project: the north star is fixed, but the route is discovered one milestone at a time. Each milestone is a hypothesis, run as a small experiment with explicit success and kill criteria β and almost everything downstream can change once the first real training run is evaluated.
- The first milestone is concept quality: find a training objective and architecture that form a quality concept bottleneck β semantically meaningful, geometrically diverse, and actually used by the decoder.
- The next milestone is recursive reasoning over those concepts: the broad idea is to refine concepts with a shared block applied multiple times, but the details are deliberately open and will be shaped by what the first experiments reveal.
The day-to-day focus moves quickly, so it is not pinned here. The living source of truth is the agenda and the active experiment spec:
- Live agenda:
docs/1_Strategy_and_Plans/agenda.md - Active experiment:
docs/experiments_specs/ - Run ledger:
docs/2_Experiments_Registry/master_experiment_log.md
The project has already run many small-scale experiments. The useful lesson is not "the old model worked" β it is the opposite: easy objectives exposed exactly where concept bottlenecks fail.
- Self-reconstruction is the wrong pressure. When the encoder sees the same content the decoder reconstructs, the system learns positional or surface shortcuts instead of semantics.
- Diversity alone is not meaning. Some regularizers raise effective rank while damaging downstream semantics.
- Parallel decoders are not enough. Position-only reconstruction can train a useful probe, but it does not prove the model can generate.
- Bidirectional token β concept interaction matters. The token side and concept side must evolve together; static token embeddings leave the bottleneck too weak.
- Concept collapse is measurable. Effective rank, pairwise concept cosine, STS-B, and concept-ablation loss are all tracked, because loss curves alone can be misleading.
- Operating regime matters as much as interface design. On frozen Gemma-3-1B, short-context plain CE kept healthy concept geometry but near-zero beyond-local causal use; the same shared-depth workspace under longer documents and more compute became load-bearing. Architecture and data/length pressure have to be judged together.
The method is deliberately incremental: one experiment, one changed variable, explicit success and kill criteria. Past runs are treated as evidence that improved understanding β not as wins or losses.
The path is genuinely open, but the direction is stable:
| Stage | Goal |
|---|---|
| 1 Β· Concept quality | Concepts that are semantically rich, geometrically non-collapsed, and useful for generation. |
| 2 Β· Concept-conditioned generation | Move from representation probes to real AR or diffusion generation from concepts. |
| 3 Β· Instruction following | Encode instructions into concepts and generate useful responses through the bottleneck. |
| 4 Β· Recursive latent reasoning | Apply a shared reasoning block repeatedly over concepts, with more refinement steps available at inference time. |
| 5 Β· Long context | Scale length through concept compression, memory, and curricula β with 10M tokens as the north-star target. |
| 6 Β· Audio-native reasoning | Map speech into concepts, reason without mandatory text round-tripping, and decode back to speech. |
The eventual audio path:
User speech β audio adapter β concept space β recursive refinement β talker / decoder β spoken response
Text is the first proving ground because it gives fast iteration, mature datasets, and clear evaluation. The larger ambition is a modality-agnostic concept space.
.
βββ nn/ # Core PyTorch model components
β βββ concept_encoder.py # Shared config, encoder, BiXT-style blocks
β βββ concept_encoder_perceiver.py
β βββ concept_encoder_weighted.py # Historical-checkpoint evaluation support
β βββ concept_losses.py
β βββ loss_manager.py
βββ training/ # Training entrypoints and shared utilities
β βββ train_concept_pretraining.py
β βββ train_perceiver_denoise.py # Temporary compatibility wrapper
β βββ utils_training.py
βββ parked/ # Weighted-MLM trainer + revivable diffusion snapshots
βββ evaluation/ # GLUE, STS-B, PAWS/SICK, checkpoint routing
βββ analysis/ # Concept-rank and geometry analysis
βββ scripts/ # Generic runner, experiment wrappers, pipelines, and utilities
βββ docs/
β βββ 1_Strategy_and_Plans/ # Current agenda and long-term vision
β βββ experiments/ # Frozen experiment specs and plans
β βββ 2_Experiments_Registry/ # Run ledger and reports
β βββ 3_Evaluations_and_Baselines/
β βββ 4_Research_Notes/ # Internal notes: dated diagnoses (append-only) + undated frameworks/strategy
β βββ literature_review/ # External paper / publication reviews (topical, ### TL;DR per paper)
β βββ 5_Archive/ # Historical plans (not current truth)
βββ parked/ # Revivable but inactive experiment families
βββ tests/
βββ verification/
βββ CHANGELOG.md
βββ pyproject.toml # uv / PEP 621 dependencies
The foundation is shared on purpose: new experiments are config-selectable extensions over the common code, not one-off training forks.
Prerequisites
- Python 3.12
- uv for dependency and environment management
- CUDA for serious training; macOS CPU/MPS is suitable for smoke tests only
Install
git clone https://github.com/ksopyla/MrCogito.git
cd MrCogito
uv syncVerify the environment
uv run python verification/torch_test.py
uv run pytest tests/ -vMain maintained training entrypoint:
uv run python training/train_concept_pretraining.py \
--hidden_size 512 \
--num_hidden_layers 6 \
--concept_num 128Remote multi-GPU launchers live in scripts/, for example:
bash scripts/train_concept_pretraining_multigpu.shExperiment protocols use thin wrappers such as scripts/launch_e05.sh and
scripts/launch_e10.sh; see scripts/README.md for launcher roles. The historical
train_perceiver_denoise_multigpu.sh path remains a temporary compatibility wrapper.
Evaluate a checkpoint:
uv run python evaluation/evaluate_model_on_glue.py \
--model_path "Cache/Training/your_checkpoint" \
--task mrpcEvery serious run is logged to a public Weights & Biases project β anyone can follow the work live, with full hyperparameters, git commit, losses, concept metrics, and evaluation results. Human-readable conclusions live in the experiment registry.
This work sits at the intersection of several converging lines of research:
- Perceiver / Perceiver IO β cross-attention bottlenecks as a general encode-reason-decode skeleton.
- Flamingo, BLIP-2 β learned query bottlenecks that condition strong decoders.
- SODA β representation learning through bottleneck generation, not trivial self-reconstruction.
- BiXT β bidirectional token β concept interaction.
- Large Concept Models, SONAR-LLM β generation and reasoning above the token level.
- Coconut & latent chain-of-thought β reasoning in continuous hidden states instead of only text.
- Recurrent / recursive transformers β test-time compute scaling through repeated refinement.
- Latent multi-agent communication β the future direction where concept vectors become the channel between cooperating models.
These are influences, not dependencies. The research question is whether a compact concept state can become the primary working memory of a generative model.
This is an active research repository β public, MIT-licensed, and intentionally transparent about negative results. The aim is not to polish a benchmark number, but to prove whether a concept bottleneck can support generation, and then reasoning.
- Long-term vision:
docs/1_Strategy_and_Plans/vision_and_goals.md - Live agenda & experiments:
docs/1_Strategy_and_Plans/agenda.mdΒ·docs/experiments_specs/ - Live experiment tracking: the open Weights & Biases project β every run, public.
- Current platforms: from-scratch concept AR / Perceiver family, plus the Gemma-3-1B + LoRA shared-depth concept workspace (
nn/backbone_concept_lm.py). Long-context Gemma training cleared a causal-use gate; semantic and reasoning follow-ups are open alongside E08 / diffusion / other designs. - Diffusion and prefix diffusion remain parked and revivable. The historical weighted-MLM trainer is parked for reproduction while its model stays loadable for evaluation. The superseded recursive-MLM fork is preserved in git history; recurrent-memory research continues on the shared concept-pretraining and backbone-concept foundations.
@misc{mrcogito2025,
title = {MrCogito: Concept Bottleneck Encoder for Long-Context Reasoning},
author = {Sopyla, Krzysztof},
year = {2025},
url = {https://github.com/ksopyla/MrCogito}
}Footnotes
-
Jaegle et al., Perceiver: General Perception with Iterative Attention, ICML 2021 β https://arxiv.org/abs/2103.03206. See also Perceiver IO, ICLR 2022 β https://arxiv.org/abs/2107.14795. β©
-
Kuratov et al., Cramming 1568 Tokens into a Single Vector and Back Again, ACL 2025 β https://arxiv.org/abs/2502.13063. β©
-
Ouro team, Ouro: Looped Language Models, 2025 β https://arxiv.org/abs/2510.25741. β©
-
Tencent & Tsinghua, CALM: Continuous Autoregressive Language Models, 2025 β https://arxiv.org/abs/2510.27688. β©
-
Geiping et al., Scaling up Test-Time Compute with Latent Reasoning (Huginn), NeurIPS 2025 β https://arxiv.org/abs/2502.05171. β©
-
Alayrac et al., Flamingo: a Visual Language Model for Few-Shot Learning, NeurIPS 2022 β https://arxiv.org/abs/2204.14198. β©
-
LCM team (Meta), Large Concept Models: Language Modeling in a Sentence Representation Space, 2024 β https://arxiv.org/abs/2412.08821. β©
-
Princeton, UIUC & Stanford, LatentMAS: Multi-Agent Collaboration in Latent Space, 2025 β https://arxiv.org/abs/2511.20639. β©
-
Zhejiang & Alibaba, Interlat: Inter-Agent Communication in Latent Space, ACL 2026 β https://arxiv.org/abs/2511.09149. β©