Named from 貯める (tameru) — Japanese for "to save, store up, or accumulate."
Tameru is a query-aware, deterministic, purely extractive context compaction engine for autonomous LLM agents. v1.3 adds typed-edge dependency restoration, a SelfCompact-style timing gate, plan-aware multi-query, qualifier-aware structured trimming, a derived-objective mode for empty queries, and a context-calibrated compression ceiling — all without external LLM calls, GPU dependencies, or runtime dependencies in the deterministic extract path (the opt-insummarisetier is the only exception, and it falls back toextracton any LLM failure).
- 💡 Why Tameru? (The Problem with Abstractive Summarization)
- 🎯 Core Design Tenets
- 🏗️ Architectural Blueprint & Data Flow
- 📐 Mathematical Formulation & Scoring Engine
- 🛡️ 6-Stage Defensive Pipeline
- 🧩 Supported Modalities & Preprocessing Engine
- 🌍 Unicode, Direction & Vertical Text
- 🏭 Industrial Limits & Format Contracts
- 📊 Head-to-Head Competitive Benchmark
- 🧪 The Production QA Battery (v3)
- 🔄 Reversible Compaction & ARC Citations (CCR Store)
- 🔌 Hermes Context-Engine Plugin Integration
- 🚀 Quickstart & Installation
- 💻 Python API & Agent Integration
- 📜 Changelog
- 🙏 Credits & Research Lineage
- 🤝 Community & Support
- 📄 License
Modern autonomous coding and research agents rapidly saturate their context windows (32k–128k+) with bulky tool outputs: git log, npm test traces, JSON API responses, database schemas, and multi-file diffs.
Traditional context reduction approaches fail in mission-critical agent workflows:
-
Abstractive LLM Summarizers (e.g., secondary LLM calls):
- Slow & Expensive: Adds 2–5 seconds of latency and doubles API costs per turn.
- Hallucinatory & Lossy: Paraphrases lose critical hex hashes, line numbers, variable names, and exact error traces.
- Context Window Pressure: Compressing a 100k context requires an auxiliary model with at least a 100k window.
-
Naive Sliding Windows / Head-Tail Truncation:
- Lost in the Middle: Drops critical intermediate facts, configurations, and multi-hop clues located in the middle turns.
Tameru operates purely extractively. Instead of generating new summary prose, it scores, filters, and retains original, verbatim text blocks while eliminating noise, repetition, and dead weight.
| Raw context (500 KB) | ⚡ Tameru | Compacted extractive view (12 KB) | |
|---|---|---|---|
| Contents | Build logs (80k) · git history (40k) · JSON APIs (120k) · diffs (60k) · … | Query-aware keep/drop scoring — no generated text | Exact error trace (lines 42–45, verbatim) · query entity references · JSON/YAML skeleton |
| Cost | full tokens | <1.2 s at 500 KB, $0 | 97.6% token reduction |
| Guarantees | — | deterministic, fail-open, reversible | 0 hallucinations · 100% byte-identical output |
Tameru adheres to a strict engineering contract:
- Query-Aware Precision: Scored against the user's active intent using sublinear BM25-IDF weighting, alphanumeric identifier extraction, and non-Latin script tokenization.
-
Strictly Deterministic: Given the exact same context string and query, Tameru produces byte-identical output every single time, verified across arbitrary
PYTHONHASHSEEDvalues. - Fail-Open Contract: If query signal is ambiguous, low-confidence, or zero-overlap, Tameru safely returns the original context rather than risking data corruption.
-
Reversible by Design: Omitted sections are replaced with content-addressed ARC citation anchors (
[A hash] "head"..."tail"), backed by a local Content-Centric Retrieval (CCR) store. -
Zero External Dependencies: Pure Python standard library (
re,json,difflib,math,typing). No PyTorch, no HuggingFace, no network required for core operations. -
Multi-Hop & Temporal Integrity: Solves graph closure (
$A \to B \to C$ ) across isolated tool outputs and prunes superseded obsolete statements. - Logical-Order Unicode Safety: Profiles LTR, RTL, mixed and vertical-source text without visually reordering or rewriting caller bytes.
- Bounded Industrial Work: Enforces input, line, record, field, profile and bidi-control limits before expensive transforms.
One pass, six stages. Every stage is deterministic; every unsafe condition exits through the same fail-open door — the caller's exact bytes come back.
| # | Stage | Input → output | Key work | On failure |
|---|---|---|---|---|
| 1 | Industrial preflight | raw text + query → profile | size/line/record/bidi limits; script, direction & format detection | hard limits → return original |
| 2 | Exact format adapters | profiled text → framed blocks | JSON/NDJSON, CSV/TSV, Markdown, YAML, XML/HTML, SQL, INI, OCR — line-record and subtree framing | malformed structure → generic scorer fallback |
| 3 | Multi-signal scoring | blocks → scored blocks | sublinear BM25-IDF, entity-density anchoring, trust & injection flags, temporal supersession | ambiguous/low-signal → fail-open later |
| 4 | Graph closure & rescue | scored blocks → keep-set | multi-hop BFS on rare shared terms; dependency closure restores qualifier/definition blocks a kept block needs | counterfactual overlap → ambiguous-failopen |
| 5 | Reversible emission | keep-set → compacted text | ARC citations [A hash] "head"…"tail", lost-in-middle reorder, CCR write (skipped on secrets) |
CCR write fails → pointer removed, text still valid |
| 6 | Self-check verifier | output → receipt | entity/keyword/critical-line recall; risk floor; selection path + sufficiency_restored reporting |
low recall → risk raised, never silently |
Cross-cutting gates sit beside the pipeline rather than inside it: the
timing gate (transcript adapter) suppresses the Tameru prune pass on
pending tool calls and stuck loops — in Hermes it gates
apply_extractive_tool_prune only, not any other compaction the host may
run; the secrets screen can veto CCR persistence; the recursion
guard refuses to compact Tameru's own output.
Each discrete text block
To prevent repetitive term spamming from dominating relevance:
Where
For multi-hop reasoning chains (
The selection floor
- Pre-Flight Gate: Evaluates
inspect_compressibility(). If repetition ratio is low or input is already compact, passes through unchanged. - Log & Structural Preprocessing: Performs count-preserving template deduplication (
[A-N] exemplar), ANSI stripping, and JSON array flattening. - Multi-Signal Scoring & Graph Closure: Computes lexical, entity, structural, and graph-expansion scores.
- Defensive Invariant Checks: Verifies that active errors, stack traces, and code fences are preserved intact (
error-signal invariant). - Lost-in-the-Middle Reordering (
reorder_best): Places top-ranked evidence at the front and second-best at the tail to optimize LLM decoder attention. - Self-Check Diagnostic Verifier: Performs post-compaction validation:
| Modality | Specialized Processing | Typical Reduction |
|---|---|---|
| Terminal & Build Logs | Count-preserving deduplication, ANSI stripping, stack frame collapse | 90–96% |
| JSON Schemas & APIs | Key-preserving array crushing, bounded nesting, exact selector matching | 75–88% |
| NDJSON / JSONL | Per-record validation and exact raw-line selection | 80–99% |
| CSV / TSV | Quote-aware record framing, header retention, embedded-newline safety | 70–95% |
| Markdown | Heading ancestry and atomic fenced-code sections | 60–90% |
| YAML | Parent-key plus matching-subtree retention; conservative decline on ambiguous prose | 60–90% |
| XML / HTML | Exact line-oriented child selection with preserved wrappers; unsafe multiline shapes decline | 50–90% |
| SQL | Quote/comment/dollar-string-aware statement selection | 60–95% |
| INI / TOML-style sections | Complete matching-section retention | 60–95% |
| Git Dumps & Commit Logs | Commit hash retention, author/subject line-record filtering | 80–92% |
| Code Diffs & Patches | Structural fence invariant, changed-line hunk isolation | 65–80% |
| Multilingual / Mixed Direction | Grapheme-safe logical-order matching across 20 script/language families | 60–95% |
| Vertical OCR Columns | Blank-column framing with logical-source-order matching | 50–90% |
Tameru never applies visual bidi reordering and never rewrites output into a different normalisation form. Original substrings remain the source of truth. A separate NFKC/case-folded matching shadow is used only for search.
- Extended grapheme tailoring keeps combining marks, variation selectors, emoji modifiers, flags, Indic virama sequences and ZWJ/ZWNJ sequences atomic.
- Direction profiles report
ltr,rtl,mixed, orneutralusing Unicode bidi classes. Explicit controls and overrides are counted separately. - Script-aware query units cover Arabic, Hebrew, Persian/Urdu, Devanagari, Bengali, Tamil, Telugu, Thai, Lao, Khmer, Myanmar, Han, Kana, Hangul, Greek, Cyrillic, Armenian, Georgian, Ethiopic and Mongolian families.
- Space-free scripts use bounded grapheme n-grams; spaced scripts use logical word runs. ASCII keeps the original compiled-regex fast path.
- CSS
writing-modeand one/two-grapheme OCR columns are metadata hints only; they never reverse or rotate source text.
See docs/industrial-compaction-v1.2.md
for the standards, tailoring decisions and invariants.
Default limits are deliberately conservative and configurable per call:
| Limit | Default |
|---|---|
| Input characters | 8,000,000 |
| Logical lines | 250,000 |
| Structured records | 100,000 |
| Characters per record | 1,000,000 |
| Fields per structured record | 4,096 |
| Unicode profile sample | 16,384 chars at each edge |
| Bidi controls | 10,000 |
| Bidi overrides | 128 |
If a hard limit, malformed surrogate, parser invariant, or structure check fails, Tameru returns the caller's exact text before CCR writes or lossy work. Adapters only activate when they produce a shorter structurally valid exact subset; otherwise the v1.1 scorer remains the fallback.
from tameru import IndustrialLimits
from tameru.compress_context import compress_context
limits = IndustrialLimits(
max_input_chars=4_000_000,
max_records=50_000,
max_bidi_overrides=16,
)
result = compress_context(context_data, query, limits=limits)
profile = result.receipt["industrial"]["profile"]Equivalent CLI controls are available as --max-input-chars, --max-lines,
--max-records, --max-record-chars, --max-fields,
--max-profile-chars, --max-bidi-controls, and --max-bidi-overrides.
Tested across 17 standardized production-QA fixtures containing multi-hop reasoning, temporal overrides, noisy logs, and structured sample text:
| Compaction System | Gold Fact Recall | Latency (avg) | Cost / 1k Ops | Deterministic | Dependencies |
|---|---|---|---|---|---|
| ⚡ Tameru (v1.3.0) | 17 / 17 (100%) | 4–956 ms observed; <1.2 s at 500 KB | $0.00 | 100% Yes | Python stdlib |
| BM25 / Vector RAG Baseline | 12 / 17 (70.6%) | ~400 ms | $0.02–$0.05 | No | Vector DB + Embeddings |
| Abstractive LLM Summarizer | 7 / 17 (41.2%) | ~2,500 ms | $1.50–$3.00 | No (Stochastic) | Auxiliary LLM API |
| Uncompressed Baseline | 17 / 17 (100%) | 0 ms | Full Tokens | Yes | None |
Same 12-case QA corpus, measured via
benchmarks/jev_comparison.py. Arms: real
TypeSafe JEV via LCC (typed
noul keep-probability questions — the real System One protocol), its
local Laya decision model (laya-multilingual, CPU) and mechanical
fallback, headroom-ai's Rust
TextCrusher, and a stdlib TF-IDF retrieval baseline:
| Metric | Tameru (v1.3.0) | lcc → JEV | lcc → Laya | lcc mechanical | headroom | TF-IDF |
|---|---|---|---|---|---|---|
| Gold retention | 12/12 | 11/11 † | 12/12 | 12/12 | 10/12 | 11/12 |
| Forbidden distractors kept | 0 | 5 | 5 | 5 | 4 | 5 |
| Deterministic | ✅ byte-identical | ❌ | ✅ | ✅ | ✅ | ✅ |
| Median latency | 10.75 ms | 591 ms † | 8,675 ms | 4.5 ms | 0.5 ms | <1 ms |
| Mean savings | 81.0% | 63.6% | 10.7% | 48.4% | 49.6% | 55.2% |
| Cost / runs local | $0 / ✅ | API-priced / ❌ | $0 / ✅ | $0 / ✅ | $0 / ✅ | $0 / ✅ |
† JEV skipped the 4,000-block perf case (API cost); Laya ran it — ~56 min on CPU.
The decisive gap isn't relevance — most judges kept the gold. It's
admissibility: every non-Tameru arm kept planted distractors,
including a stale superseded config and EXCLUDED-HOST inside a block
labeled "UNTRUSTED SAMPLE". A keep-score — JEV probability, Laya decision,
BM25, or TF-IDF — has no trust, supersession, or exclusion model; Tameru
encodes all three. Full table, per-case detail, and methodology:
benchmarks/COMPARISON.md. Ecosystem +
paper digest: docs/RESEARCH.md.
The tests/ directory contains an exhaustive production-QA suite:
test_arabic_and_sql_sink.py: Non-Latin script preservation and SQL tabular sink compaction.test_adjacent_v010.py: Compaction audit logging, pinned sink regions, and versioned context receipts.test_hardening_pass_v09.py: NTK template deduplication, progress-bar stripping, and stack frame collapse.test_semantic_gates.py: Counterfactual distractor masking and graph closure validation.test_production_qa_v3.py: Cross-seed determinism validation (PYTHONHASHSEED=0vsPYTHONHASHSEED=1).
Run the full battery:
python -m unittest discover -s testsCurrent v1.3.0 release verification:
- 326 passed, 14 skipped with pytest.
- 15/15 production-QA battery cases passed (
benchmarks/run_battery.py), including Japanese, Chinese and Arabic needle cases with zero forbidden-distractor leaks. benchmarks/threshold_sweep.py: everybudget_ratiooperating point from 0.1–0.7 keeps all gold with 0 leaks — a clean 77%→25% savings frontier.- Six-arm live comparison (Tameru / JEV / Laya / mechanical / headroom / TF-IDF): only arm with 12/12 gold + 0 leaks + deterministic output — table below.
When Tameru drops content blocks, it emits content-addressed ARC Citations:
[A 7f8a9b] "2026-08-25 02:00:00 [INFO] Worker started"...[48 lines omitted]..."2026-08-25 02:05:00 [INFO] Complete"
If an agent subsequently determines that it needs the omitted content, it can query the local Content-Centric Retrieval (CCR) store:
from tameru.compress_context import retrieve
full_context = retrieve("7f8a9b2c4d6e8f0011223344")CCR is opt-in per integration path. The standalone API keeps its reversible
default, while the Hermes live tool-pruning adapter disables CCR because it has
no retrieval path and tool payloads may contain secrets. Expired valid records
are swept whenever a new CCR record is written; operators can also call
sweep_ccr_cache() directly. CCR directories and records are created with
owner-only permissions where the platform supports POSIX modes.
Tameru integrates into Hermes Agent as a pluggable context engine — no Hermes source changes:
- Standalone plugin —
0xWhiteMage/hermes-tameru-plugin: copy into~/.hermes/plugins/tameru/orpip installit (auto-discovered via thehermes_agent.pluginsentry point). - Activate in
config.yaml:context: engine: tameru
pip install -e .pip install -e '.[semantic]'To equip an autonomous agent with Tameru, paste this prompt into your agent's system prompt or session:
Please install and enable the Tameru compaction system for high-performance context management:
1. Clone or vendor tameru into your environment.
2. In your context configuration, select 'tameru' as the active context engine.
3. Use tameru.compress_context(context, query) before passing large tool outputs to your context window.
from tameru.compress_context import compress_context
context_data = (
"2026-08-25 02:00:00 [INFO] System boot sequence started.\n"
+ ("2026-08-25 02:01:00 [INFO] Routine worker heartbeat.\n" * 500)
+ "2026-08-25 02:14:00 [ERROR] DB connection failed on host db-prod-01: Port 5432 unreachable.\n"
)
query = "Why did the database connection fail?"
result = compress_context(context_data, query)
print(f"Compressed Output:\n{result.compressed_text}\n")
print(f"Token Reduction: {result.tokens_saved_pct:.1f}%")
print(f"Fail-Open Triggered: {result.fail_open}")
print(f"Confidence Diagnostic: {result.verifier}")strategy="summarise" remains fail-open and can be configured per call with
summary_endpoint, summary_models, and summary_timeout. Equivalent
environment variables are TAMERU_SUMMARY_ENDPOINT,
TAMERU_SUMMARY_MODELS (comma-separated), and TAMERU_SUMMARY_TIMEOUT
(seconds). The timeout is one total retry budget across all candidate models,
not a fresh timeout per model.
Summary endpoints are restricted to loopback by default. Set
TAMERU_SUMMARY_ALLOW_REMOTE=1 or pass summary_allow_remote=True only when
the integration has explicitly approved sending context to that endpoint.
When using decision_cache, first-sight keep/drop outcomes are replayed on
later turns. The cache preserves the oldest prefix and is bounded to 4,096
block decisions. Prefix stability intentionally outranks a later fixed-mode
budget_ratio; if frozen keeps exceed that ratio, the engine preserves the
frozen prefix and reports freeze cache capacity reached once the cache can no
longer learn new blocks.
Additional caller controls, inspired by transcript-level pruners:
pin_recent=Nunconditionally keeps the opening block and the newest N blocks — a working-context floor that no selector, freeze, supersession, or trust path may evict.min_savings_ratio=R(default0.10) fails open when achieved savings fall belowR, so callers can decline a compaction that did not shrink enough;0disables the floor.degraded_view=Truerescues input size-limit breaches (max_input_chars,max_lines): instead of failing open, each oversized block is scored on a bounded head+tail view while emitted output stays byte-exact. Safety limits (bidi controls, malformed surrogates, oversize query) remain hard failures, and work stays bounded by a 4× grace factor over the configured limits. The receipt reportsdegraded_view: trueand the risk floor becomesmedium. CLI:--degraded-view.
Further ecosystem-research hardening (deterministic, always-on unless noted):
- Secrets screen: if the input contains probable credentials (private
key blocks,
AKIA…, GitHub/Slack/sk-tokens, JWTs, long quotedpassword/api_key-style assignments), the CCR store is skipped — archival would persist secrets to disk at rest. The result reportsccr: None, emits no[CC-Retrieve:]marker, and the reason list notesccr skipped: secret material detected. - Recursion guard: input already carrying Tameru output markers
(
<compressed_context,[CC-Retrieve:) returns unchanged withpolicy_name="local-noop-recursion"— compacting compacted output would nest wrappers and could drop the pointer that makes earlier drops recoverable. - Compressible-subset budgeting: in
mode="fixed", pinned blocks are immovable and no longer consume thebudget_ratiothey outrank; the ratio governs only compressible tokens. With no pins, budgeting is unchanged. - CCR recall tooling:
list_ccr(ccr_dir, offset=, limit=)lists live records newest-first (hash, stored_at, ttl, chars, preview), andretrieve(hash, offset=, limit=)paginates large originals. - Selection transparency: every receipt reports
selection— which selector path decided the keep-set (needle,floor,floor-saturated,line-records,fixed, or a*-failopenvariant), so silent degradation is visible to callers. strategy="auto"runs the progressive ladder:extractfirst, then escalate tosummariseonly when extraction fails open (ambiguity, saturated floors,min_savings_ratioundershoot).summarisestill falls back toextracton any LLM failure, so the worst case is one extra deterministic pass. CLI:--strategy auto.
Second research round (v1.3.0 — lcc internals, SelfCompact, PAACE, TPC, Compactor):
- Dependency closure (sufficiency restore): after selection, a dropped
block that shares a rare term with a kept block AND carries a qualifier
cue (
except,unless,only,however,until,if, …) or a definition cue (is defined as,refers to,namely, …) is restored — cutting it would invert the meaning of what survives (e.g. a kept "throughput is nominal" whose dropped sibling says "except during maintenance"). Capped at 8 restorations, never resurrects trust-risk or frozen-drop blocks, supersession can still evict stale restorations, andmode="fixed"restores only within the caller's hard budget. Receipts reportsufficiency_restored: [block ids]. - Qualifier-aware trim refusal:
_crush_valueno longer truncates a long JSON string whose cut tail carries a qualifier cue — a longer safe value beats a shorter misleading one. - Timing gate (transcript adapter):
tameru.transcript.trajectory_gatesuppresses the Tameru prune pass mid-derivation (pending tool calls) and on stuck loops (the last 3 assistant turns issued identical calls — diagnose, don't erase evidence). On by default viaapply_extractive_tool_prune(..., timing_gate=True); it only ever suppresses that prune step — downstream host compaction is unaffected. - Plan-aware multi-query:
compress_context(ctx, [q1, q2, ...])scores blocks against the union of current + planned tasks. - Derived objective (
derive_query=True, CLI--derive-query): an empty/generic query normally fails open; with derivation the engine infers a conservative task descriptor from the document's own recurring rare terms (deterministic frequency statistics, bounded sampling). Receipts markquery_source: "derived"+derived_terms, the risk floor becomesmedium, and documents with no stable term structure still fail open. - Compression ceiling:
inspect_compressibility()now reportsguaranteed_savings_pct+ceiling_class(dedupe-heavy/moderate/sparse) — the share removable by pure dedupe before any relevance judgement, so callers can see how much headroom a given context actually has. benchmarks/threshold_sweep.py: replays the QA corpus acrossbudget_ratiovalues and reports the gold/leak/savings frontier — operating points chosen on evidence, not inherited.- Multilingual coverage:
zh_needle+ar_needlebattery cases — compression tuned on English silently fails other scripts.
Tameru ships harness-agnostic. Three integration paths, cheapest first — see harnesses/README.md for the full contract:
- Python API —
pip install tameru-compaction-system, then callcompress_context(text, query)for raw payloads ortameru.transcript.apply_extractive_tool_prune(messages)for OpenAI-stylerole/contentconversation lists (Hermes, OpenCode, Codex, and most agent loops share that shape). - CLI —
tameru-compressreads a file or stdin, writes compacted text, and emits JSON stats with--stats; the universal shell-out for non-Python harnesses. - Vendored plugin — copy
src/tameru/*.pyinto the plugin dir and keep it current withpython scripts/sync_to_harness.py <plugin_dir> --manifest plugin.yaml. Upstream owns every module except__init__.py; the harness owns registration glue and the manifest.
All notable changes, version milestones, and migration notes are tracked in CHANGELOG.md.
Highlights:
- v1.3.0: Typed-edge dependency restoration, transcript timing gate, plan-aware multi-query, qualifier-aware trim refusal,
derive_query, compression-ceiling reporting, threshold-sweep tool, multilingual battery. - v1.2.1: JEV-inspired hardening —
pin_recent,min_savings_ratio,degraded_view, secrets screen, recursion guard, compressible-subset budgeting, CCR recall tooling, selection receipts,strategy="auto". - v1.2.0: Bounded industrial preflight, logical-order Unicode across 20 script/language families, ten exact format adapters, deterministic receipt hashes and large-input SLOs.
- v1.1.1: Comprehensive factual-retention, fail-open, cache-progression, public-metrics, and current-Hermes integration hardening.
- v1.1.0: Production QA hardening, Docker progress recognition, CCR security & expiry sweep, bounded decision caching, loopback summary boundary.
- v0.10.0: Compaction audit logs, pinned sink regions, versioned context receipts.
- v0.9.0: NTK layer-1 log dedup, error invariants, progress-bar stripping, lost-in-the-middle reordering.
- v0.8.0: Opt-in semantic tier with bi-encoders and cross-encoders.
- v0.1.0 - v0.7.0: Core extractive engine, graph closure, temporal supersession, and tabular crushers.
Tameru's design was inspired and sharpened by studying remarkable open-source projects, papers, and ideas:
- NTK / Neural Token Killer (Rust, MIT) — Template deduplication, error-signal invariant, progress-bar stripping, and stack frame collapse.
- caveman-compression (MIT) — Executable test specification discipline and factual preservation validation.
- leanctx (MIT) — Loss-tolerance routing and structural verbatim code invariants.
- TwoTrim (Apache-2.0) — Lost-in-the-middle mitigation and attention edge anchoring.
- clipforge-PAKT (MIT) — Pre-flight compressibility inspection.
- fast-jev-compaction (MIT) — Positional pinning of edge turns, caller-side savings-worthiness gates, and judging on a degraded state while emitting exact spans.
- dsh-argp — Compressible-subset retention budgeting: immovable atoms are excluded from the ratio's base.
- dsh-jev-prune — Never degrade silently: report which selector path decided, and hard-exclude mutating actions by rule rather than by score.
- dsh-jev-pre-compaction — Scan for secrets before archiving originals; lookahead pressure bands.
- pi-lcm — Persistent append-only memory with self-serve grep/describe/expand recall tools.
- lcc / Local Context Compiler (MIT) — Typed context-graph closure, post-drop sufficiency verification, qualifier-aware safe trimming, and a well-built JEV/local-model client used as our benchmark harness.
- headroom — Closest production analog: extractive crushers + reversible CCR + cache alignment; its CJK token-pricing fix prompted our (passing) audit.
- Waxmell114514/jev-compaction — Score-only reversible compaction and the threshold-replay tuning methodology behind
benchmarks/threshold_sweep.py. - SelfCompact — The closed-unit/progress/not-stuck compaction-timing rubric, ported as
tameru.transcript.trajectory_gate. - Claude Code compaction teardown (anneheartrecord/claude-code-docs) — Progressive micro→session→full tiers, recursion guards on compaction agents, and circuit breakers after repeated failures.
- LongLLMLingua / LLMLingua-2 (Microsoft, ACL 2024) — Question-aware compression and the preserve/discard token-classification formulation.
- Lost in the Middle (Liu et al., 2023) — Positional attention decay in decoder transformers.
- NoLiMa (2025) — Long-context lexical distractor blind spots.
- Context Rot (Chroma, 2025) — The empirical case for deliberate compression: models degrade as input grows even on simple tasks.
- Provence (2025) — "Zero sentences kept" as a legal answer; pruning and ranking as one operation.
- PAACE (2025) — Plan-aware selection for next-k tasks, ported as multi-query union scoring.
- TPC (2025) — Task-descriptor objectives for question-free compression, ported deterministically as
derive_query. - Compactor (2025) — Context-calibrated compression ceilings, ported as
inspect_compressibilityguaranteed-savings reporting. - Lost in Compression (2026 audit) — Cross-lingual failure modes of English-tuned compressors; drove the multilingual battery.
- ARC Citations — Reversible content-addressed reference anchors.
Full source-by-source digest: docs/RESEARCH.md.
- Author: Benjamin Ang (@0xWhiteMage)
- Support: Buy me a coffee on Ko-fi
- Issues & Contributions: Pull requests and issue reports are welcome!
Released under the MIT License.
