Skip to content

Latest commit

Β 

History

628 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

MrCogito

A research project building models that compress long context into latent concept vectors, reason in that concept space, and decode back to text or other modalities.

Author: Krzysztof Sopyla β€” ai.ksopyla.com Β· GitHub Β· LinkedIn Project page: ai.ksopyla.com/projects/concept-encoder Experiments: wandb.ai/ksopyla/MrCogito


Vision

Today's language models reason by expanding thought into text tokens. That works, but it is an extremely narrow channel: every intermediate idea has to be serialized into a vocabulary item, attended over again, and paid for again in the context window.

This project explores a different primitive:

Compress input into dense latent concepts, refine those concepts recursively, and generate output from the refined concept state.

The end goal is a foundation-model architecture defined by four bold targets:

  • 🧠 Reasoning in latent space. The model should spend most of its compute refining continuous concept states, not just emitting visible chain-of-thought tokens.
  • πŸ“œ 10M-token context as the north star. Reached not by forcing full self-attention to scale forever, but by compressing long streams into a far smaller set of concept vectors.
  • πŸ”€ One concept substrate across modalities. Text, audio, and eventually other inputs map into the same reasoning space, with modality-specific adapters and decoders only at the edges.
  • 🀝 A path to latent agent communication. Once concept vectors are a reliable semantic interface, models can exchange concepts directly instead of talking through text. That multi-agent layer lives outside this repository β€” but it is the foundation this work is meant to enable.

The core bet is simple: if concept vectors can carry rich semantic and generative state, long-context reasoning becomes cheaper, more inspectable, and more composable than token-only reasoning.


Why This Project Exists

The frontier's answer to long context has been more of the same: better hardware, clever attention approximations, bigger clusters. I want a different bet β€” not making self-attention cheaper, but asking whether the model needs it at all for most of its reasoning.

A few older ideas I think were ahead of their time, and that I believe belong together:

  • Perceiver1 β€” a small set of latent queries can cross-attend to a huge input and decouple compute from sequence length. Better parameter and compute utilization through cross-attention β€” but used for perception, not as the seat of reasoning.
  • Embedding space is wildly underused. A single dense vector can be made to hold ~1,500 tokens of recoverable text2, yet we still spend enormous capacity on giant vocabularies and wide token embeddings. I believe most of that capacity goes to surface form, not meaning β€” the latent space is the asset we're not exploiting.
  • Reasoning can move into latent space. Looping the same weights3 and predicting continuous vectors instead of discrete tokens4 let a model "think" beyond the token channel and punch above its parameter count5.

My bet is that combining these into one end-to-end trained architecture β€” a cross-attention concept bottleneck plus a recursive reasoning core operating in latent space β€” opens a different paradigm than the scaling race. Cross-attention bottlenecks themselves aren't new (Flamingo6, Large Concept Models7), but they're treated as a means to an end, not as the primary locus of reasoning. The question I keep asking is sharper:

Can a small set of concept vectors become the model's working memory β€” the thing it actually reasons over?

If yes, the payoff compounds: O(CΒ·N) long context with C β‰ͺ N, test-time compute scaling by running more refinement steps without retraining, modality-agnostic reasoning over a shared concept space, and an internal state small enough to actually inspect. It also opens a channel the token interface can't: letting agents talk to each other in concept space instead of text β€” early work shows passing latent state between models is both faster and more accurate than exchanging tokens89. It may not work β€” it's a genuine research bet β€” but I think it's the more interesting direction, and the one the scaling race is leaving unexplored.


Architecture

The architecture follows an encode β†’ reason β†’ decode pattern, with the concept space at its center:

Input tokens / frames  (N)
        β”‚
        β–Ό
   Encoder            cross-attention compresses N inputs into C concepts   (C β‰ͺ N)
        β”‚
        β–Ό
 Reasoning core       optional recursive refinement over the concepts
        β”‚
        β–Ό
   Decoder            generates text, speech, or another modality from concepts

A small set of learned concept queries cross-attends to the full input sequence. This shifts the dominant attention pattern from O(NΒ²) token-to-token attention to O(CΒ·N) concept-to-token attention β€” and the concept count C grows far more slowly than the input length N.

Input length N Concepts C Attention savings vs. full self-attention
512 128 ~4Γ—
32K 2K ~16Γ—
1M 8K ~128Γ—
10M hierarchical north-star target β€” needs hierarchical concepts, recurrence, and memory

This is not just an efficiency trick. The bottleneck is meant to force abstraction: concepts become the model's working memory, the object reasoning operates on, and the interface decoders read from.

Two complementary platforms

The project explores the same concept-bottleneck idea in two complementary settings:

  1. From-scratch concept encoder (ConceptEncoder / Perceiver–BiXT family) β€” train the encoder, reasoning surface, and decoder end-to-end. Useful for objective and geometry studies (prefixβ†’suffix AR, windowed decode, STS-B probes).
  2. Pretrained LM + concept workspace (BackboneConceptLM) β€” freeze a strong token model (Gemma-3-1B), adapt it lightly with LoRA, and attach a small set of shared depth-recurrent concepts that read from and write into selected backbone depths as the sequence is consumed in blocks (K tokens β†’ one concept update). The backbone keeps local fluency; the concepts are meant to carry multi-block state.

The Gemma path asks a sharper mechanism question: when next-token loss can already be solved from local context, do concepts ever become causally necessary? Recent long-context training (seq 4096, long-document mix, ~1B tokens) cleared that gate β€” concept ablations hurt beyond-local positions by a large margin while geometry stayed healthy. That does not close the from-scratch path or latent-reasoning / diffusion threads; it marks one validated regime worth exploring further. Details live in the agenda and E16b run report.


How the Project Works

This is an active research project: the north star is fixed, but the route is discovered one milestone at a time. Each milestone is a hypothesis, run as a small experiment with explicit success and kill criteria β€” and almost everything downstream can change once the first real training run is evaluated.

  • The first milestone is concept quality: find a training objective and architecture that form a quality concept bottleneck β€” semantically meaningful, geometrically diverse, and actually used by the decoder.
  • The next milestone is recursive reasoning over those concepts: the broad idea is to refine concepts with a shared block applied multiple times, but the details are deliberately open and will be shaped by what the first experiments reveal.

The day-to-day focus moves quickly, so it is not pinned here. The living source of truth is the agenda and the active experiment spec:


What Has Been Learned

The project has already run many small-scale experiments. The useful lesson is not "the old model worked" β€” it is the opposite: easy objectives exposed exactly where concept bottlenecks fail.

  • Self-reconstruction is the wrong pressure. When the encoder sees the same content the decoder reconstructs, the system learns positional or surface shortcuts instead of semantics.
  • Diversity alone is not meaning. Some regularizers raise effective rank while damaging downstream semantics.
  • Parallel decoders are not enough. Position-only reconstruction can train a useful probe, but it does not prove the model can generate.
  • Bidirectional token ↔ concept interaction matters. The token side and concept side must evolve together; static token embeddings leave the bottleneck too weak.
  • Concept collapse is measurable. Effective rank, pairwise concept cosine, STS-B, and concept-ablation loss are all tracked, because loss curves alone can be misleading.
  • Operating regime matters as much as interface design. On frozen Gemma-3-1B, short-context plain CE kept healthy concept geometry but near-zero beyond-local causal use; the same shared-depth workspace under longer documents and more compute became load-bearing. Architecture and data/length pressure have to be judged together.

The method is deliberately incremental: one experiment, one changed variable, explicit success and kill criteria. Past runs are treated as evidence that improved understanding β€” not as wins or losses.


Long-Term Roadmap

The path is genuinely open, but the direction is stable:

Stage Goal
1 Β· Concept quality Concepts that are semantically rich, geometrically non-collapsed, and useful for generation.
2 Β· Concept-conditioned generation Move from representation probes to real AR or diffusion generation from concepts.
3 Β· Instruction following Encode instructions into concepts and generate useful responses through the bottleneck.
4 Β· Recursive latent reasoning Apply a shared reasoning block repeatedly over concepts, with more refinement steps available at inference time.
5 Β· Long context Scale length through concept compression, memory, and curricula β€” with 10M tokens as the north-star target.
6 Β· Audio-native reasoning Map speech into concepts, reason without mandatory text round-tripping, and decode back to speech.

The eventual audio path:

User speech  β†’  audio adapter  β†’  concept space  β†’  recursive refinement  β†’  talker / decoder  β†’  spoken response

Text is the first proving ground because it gives fast iteration, mature datasets, and clear evaluation. The larger ambition is a modality-agnostic concept space.


Repository Guide

.
β”œβ”€β”€ nn/                         # Core PyTorch model components
β”‚   β”œβ”€β”€ concept_encoder.py        # Shared config, encoder, BiXT-style blocks
β”‚   β”œβ”€β”€ concept_encoder_perceiver.py
β”‚   β”œβ”€β”€ concept_encoder_weighted.py  # Historical-checkpoint evaluation support
β”‚   β”œβ”€β”€ concept_losses.py
β”‚   └── loss_manager.py
β”œβ”€β”€ training/                   # Training entrypoints and shared utilities
β”‚   β”œβ”€β”€ train_concept_pretraining.py
β”‚   β”œβ”€β”€ train_perceiver_denoise.py  # Temporary compatibility wrapper
β”‚   └── utils_training.py
β”œβ”€β”€ parked/                     # Weighted-MLM trainer + revivable diffusion snapshots
β”œβ”€β”€ evaluation/                 # GLUE, STS-B, PAWS/SICK, checkpoint routing
β”œβ”€β”€ analysis/                   # Concept-rank and geometry analysis
β”œβ”€β”€ scripts/                    # Generic runner, experiment wrappers, pipelines, and utilities
β”œβ”€β”€ docs/
β”‚   β”œβ”€β”€ 1_Strategy_and_Plans/     # Current agenda and long-term vision
β”‚   β”œβ”€β”€ experiments/             # Frozen experiment specs and plans
β”‚   β”œβ”€β”€ 2_Experiments_Registry/   # Run ledger and reports
β”‚   β”œβ”€β”€ 3_Evaluations_and_Baselines/
β”‚   β”œβ”€β”€ 4_Research_Notes/         # Internal notes: dated diagnoses (append-only) + undated frameworks/strategy
β”‚   β”œβ”€β”€ literature_review/        # External paper / publication reviews (topical, ### TL;DR per paper)
β”‚   └── 5_Archive/               # Historical plans (not current truth)
β”œβ”€β”€ parked/                     # Revivable but inactive experiment families
β”œβ”€β”€ tests/
β”œβ”€β”€ verification/
β”œβ”€β”€ CHANGELOG.md
└── pyproject.toml              # uv / PEP 621 dependencies

The foundation is shared on purpose: new experiments are config-selectable extensions over the common code, not one-off training forks.


Setup

Prerequisites

  • Python 3.12
  • uv for dependency and environment management
  • CUDA for serious training; macOS CPU/MPS is suitable for smoke tests only

Install

git clone https://github.com/ksopyla/MrCogito.git
cd MrCogito
uv sync

Verify the environment

uv run python verification/torch_test.py
uv run pytest tests/ -v

Training and Evaluation

Main maintained training entrypoint:

uv run python training/train_concept_pretraining.py \
  --hidden_size 512 \
  --num_hidden_layers 6 \
  --concept_num 128

Remote multi-GPU launchers live in scripts/, for example:

bash scripts/train_concept_pretraining_multigpu.sh

Experiment protocols use thin wrappers such as scripts/launch_e05.sh and scripts/launch_e10.sh; see scripts/README.md for launcher roles. The historical train_perceiver_denoise_multigpu.sh path remains a temporary compatibility wrapper.

Evaluate a checkpoint:

uv run python evaluation/evaluate_model_on_glue.py \
  --model_path "Cache/Training/your_checkpoint" \
  --task mrpc

Every serious run is logged to a public Weights & Biases project β€” anyone can follow the work live, with full hyperparameters, git commit, losses, concept metrics, and evaluation results. Human-readable conclusions live in the experiment registry.


Influences

This work sits at the intersection of several converging lines of research:

  • Perceiver / Perceiver IO β€” cross-attention bottlenecks as a general encode-reason-decode skeleton.
  • Flamingo, BLIP-2 β€” learned query bottlenecks that condition strong decoders.
  • SODA β€” representation learning through bottleneck generation, not trivial self-reconstruction.
  • BiXT β€” bidirectional token ↔ concept interaction.
  • Large Concept Models, SONAR-LLM β€” generation and reasoning above the token level.
  • Coconut & latent chain-of-thought β€” reasoning in continuous hidden states instead of only text.
  • Recurrent / recursive transformers β€” test-time compute scaling through repeated refinement.
  • Latent multi-agent communication β€” the future direction where concept vectors become the channel between cooperating models.

These are influences, not dependencies. The research question is whether a compact concept state can become the primary working memory of a generative model.


Status

This is an active research repository β€” public, MIT-licensed, and intentionally transparent about negative results. The aim is not to polish a benchmark number, but to prove whether a concept bottleneck can support generation, and then reasoning.

  • Long-term vision: docs/1_Strategy_and_Plans/vision_and_goals.md
  • Live agenda & experiments: docs/1_Strategy_and_Plans/agenda.md Β· docs/experiments_specs/
  • Live experiment tracking: the open Weights & Biases project β€” every run, public.
  • Current platforms: from-scratch concept AR / Perceiver family, plus the Gemma-3-1B + LoRA shared-depth concept workspace (nn/backbone_concept_lm.py). Long-context Gemma training cleared a causal-use gate; semantic and reasoning follow-ups are open alongside E08 / diffusion / other designs.
  • Diffusion and prefix diffusion remain parked and revivable. The historical weighted-MLM trainer is parked for reproduction while its model stays loadable for evaluation. The superseded recursive-MLM fork is preserved in git history; recurrent-memory research continues on the shared concept-pretraining and backbone-concept foundations.

Citation

@misc{mrcogito2025,
  title   = {MrCogito: Concept Bottleneck Encoder for Long-Context Reasoning},
  author  = {Sopyla, Krzysztof},
  year    = {2025},
  url     = {https://github.com/ksopyla/MrCogito}
}

License

MIT License


Footnotes

  1. Jaegle et al., Perceiver: General Perception with Iterative Attention, ICML 2021 β€” https://arxiv.org/abs/2103.03206. See also Perceiver IO, ICLR 2022 β€” https://arxiv.org/abs/2107.14795. ↩

  2. Kuratov et al., Cramming 1568 Tokens into a Single Vector and Back Again, ACL 2025 β€” https://arxiv.org/abs/2502.13063. ↩

  3. Ouro team, Ouro: Looped Language Models, 2025 β€” https://arxiv.org/abs/2510.25741. ↩

  4. Tencent & Tsinghua, CALM: Continuous Autoregressive Language Models, 2025 β€” https://arxiv.org/abs/2510.27688. ↩

  5. Geiping et al., Scaling up Test-Time Compute with Latent Reasoning (Huginn), NeurIPS 2025 β€” https://arxiv.org/abs/2502.05171. ↩

  6. Alayrac et al., Flamingo: a Visual Language Model for Few-Shot Learning, NeurIPS 2022 β€” https://arxiv.org/abs/2204.14198. ↩

  7. LCM team (Meta), Large Concept Models: Language Modeling in a Sentence Representation Space, 2024 β€” https://arxiv.org/abs/2412.08821. ↩

  8. Princeton, UIUC & Stanford, LatentMAS: Multi-Agent Collaboration in Latent Space, 2025 β€” https://arxiv.org/abs/2511.20639. ↩

  9. Zhejiang & Alibaba, Interlat: Inter-Agent Communication in Latent Space, ACL 2026 β€” https://arxiv.org/abs/2511.09149. ↩

About

Concept NN architecture - research repository with experiments for latent concept forming via attention and diffusion process.

Topics

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Used by

Contributors

Languages