Skip to content

Repository files navigation

⚡ Tameru (貯める) — Industrial Extractive Context Compaction for LLM Agents

Tameru Compaction System (貯める)

Version 1.3.0 License: MIT Test Suite Production QA Large-case latency Determinism Zero Dependencies

Follow on X  •   Ko-fi

Named from 貯める (tameru) — Japanese for "to save, store up, or accumulate."
Tameru is a query-aware, deterministic, purely extractive context compaction engine for autonomous LLM agents. v1.3 adds typed-edge dependency restoration, a SelfCompact-style timing gate, plan-aware multi-query, qualifier-aware structured trimming, a derived-objective mode for empty queries, and a context-calibrated compression ceiling — all without external LLM calls, GPU dependencies, or runtime dependencies in the deterministic extract path (the opt-in summarise tier is the only exception, and it falls back to extract on any LLM failure).


📑 Table of Contents


💡 Why Tameru? (The Problem with Abstractive Summarization)

Modern autonomous coding and research agents rapidly saturate their context windows (32k–128k+) with bulky tool outputs: git log, npm test traces, JSON API responses, database schemas, and multi-file diffs.

Traditional context reduction approaches fail in mission-critical agent workflows:

  1. Abstractive LLM Summarizers (e.g., secondary LLM calls):

    • Slow & Expensive: Adds 2–5 seconds of latency and doubles API costs per turn.
    • Hallucinatory & Lossy: Paraphrases lose critical hex hashes, line numbers, variable names, and exact error traces.
    • Context Window Pressure: Compressing a 100k context requires an auxiliary model with at least a 100k window.
  2. Naive Sliding Windows / Head-Tail Truncation:

    • Lost in the Middle: Drops critical intermediate facts, configurations, and multi-hop clues located in the middle turns.

The Extractive Alternative

Tameru operates purely extractively. Instead of generating new summary prose, it scores, filters, and retains original, verbatim text blocks while eliminating noise, repetition, and dead weight.

Raw context (500 KB) ⚡ Tameru Compacted extractive view (12 KB)
Contents Build logs (80k) · git history (40k) · JSON APIs (120k) · diffs (60k) · … Query-aware keep/drop scoring — no generated text Exact error trace (lines 42–45, verbatim) · query entity references · JSON/YAML skeleton
Cost full tokens <1.2 s at 500 KB, $0 97.6% token reduction
Guarantees — deterministic, fail-open, reversible 0 hallucinations · 100% byte-identical output

🎯 Core Design Tenets

Tameru adheres to a strict engineering contract:

  • Query-Aware Precision: Scored against the user's active intent using sublinear BM25-IDF weighting, alphanumeric identifier extraction, and non-Latin script tokenization.
  • Strictly Deterministic: Given the exact same context string and query, Tameru produces byte-identical output every single time, verified across arbitrary PYTHONHASHSEED values.
  • Fail-Open Contract: If query signal is ambiguous, low-confidence, or zero-overlap, Tameru safely returns the original context rather than risking data corruption.
  • Reversible by Design: Omitted sections are replaced with content-addressed ARC citation anchors ([A hash] "head"..."tail"), backed by a local Content-Centric Retrieval (CCR) store.
  • Zero External Dependencies: Pure Python standard library (re, json, difflib, math, typing). No PyTorch, no HuggingFace, no network required for core operations.
  • Multi-Hop & Temporal Integrity: Solves graph closure ($A \to B \to C$) across isolated tool outputs and prunes superseded obsolete statements.
  • Logical-Order Unicode Safety: Profiles LTR, RTL, mixed and vertical-source text without visually reordering or rewriting caller bytes.
  • Bounded Industrial Work: Enforces input, line, record, field, profile and bidi-control limits before expensive transforms.

🏗️ Architectural Blueprint & Data Flow

One pass, six stages. Every stage is deterministic; every unsafe condition exits through the same fail-open door — the caller's exact bytes come back.

# Stage Input → output Key work On failure
1 Industrial preflight raw text + query → profile size/line/record/bidi limits; script, direction & format detection hard limits → return original
2 Exact format adapters profiled text → framed blocks JSON/NDJSON, CSV/TSV, Markdown, YAML, XML/HTML, SQL, INI, OCR — line-record and subtree framing malformed structure → generic scorer fallback
3 Multi-signal scoring blocks → scored blocks sublinear BM25-IDF, entity-density anchoring, trust & injection flags, temporal supersession ambiguous/low-signal → fail-open later
4 Graph closure & rescue scored blocks → keep-set multi-hop BFS on rare shared terms; dependency closure restores qualifier/definition blocks a kept block needs counterfactual overlap → ambiguous-failopen
5 Reversible emission keep-set → compacted text ARC citations [A hash] "head"…"tail", lost-in-middle reorder, CCR write (skipped on secrets) CCR write fails → pointer removed, text still valid
6 Self-check verifier output → receipt entity/keyword/critical-line recall; risk floor; selection path + sufficiency_restored reporting low recall → risk raised, never silently

Cross-cutting gates sit beside the pipeline rather than inside it: the timing gate (transcript adapter) suppresses the Tameru prune pass on pending tool calls and stuck loops — in Hermes it gates apply_extractive_tool_prune only, not any other compaction the host may run; the secrets screen can veto CCR persistence; the recursion guard refuses to compact Tameru's own output.


📐 Mathematical Formulation & Scoring Engine

Each discrete text block $b_i \in B$ is evaluated through a composite scoring function balancing lexical overlap, entity density, structural importance, and recency:

$$\mathcal{S}(b_i \mid q) = \mathcal{S}_{\text{lex}}(b_i, q) + \mathcal{S}_{\text{ent}}(b_i, q) + \mathcal{S}_{\text{struct}}(b_i) + \mathcal{S}_{\text{recency}}(b_i) - \mathcal{P}_{\text{trust}}(b_i) - \mathcal{P}_{\text{stale}}(b_i)$$

1. Sublinear Lexical BM25-IDF Overlap

To prevent repetitive term spamming from dominating relevance:

$$\mathcal{S}_{\text{lex}}(b_i, q) = \sum_{t \in b_i \cap q} \text{IDF}(t) \cdot \left(1 + \ln(1 + \text{tf}(t, b_i))\right)$$

Where $\text{IDF}(t) = \ln\left(1 + \frac{|B| - n(t) + 0.5}{n(t) + 0.5}\right)$.

2. Multi-Hop Graph Closure (BFS Bridge Expansion)

For multi-hop reasoning chains ($A \to B \to C$), if an evidence block $b_j$ shares a bridge entity $e$ with a high-scoring block $b_k$, its score is boosted proportional to graph distance:

$$\mathcal{S}_{\text{graph}}(b_j) = \max_{b_k \in B_{\text{kept}}} \left( \mathcal{S}(b_k \mid q) \cdot \gamma^{\text{hop-distance}(b_j, b_k)} \right), \quad \gamma = 0.65$$

3. Adaptive Selection Threshold

The selection floor $\theta_{\text{adaptive}}$ dynamically scales based on candidate distribution:

$$\theta_{\text{adaptive}} = \max\left(2.2, \min\left(11.5, 0.38 \cdot \max_{b \in B} \mathcal{S}(b \mid q)\right)\right)$$


🛡️ 6-Stage Defensive Pipeline

  1. Pre-Flight Gate: Evaluates inspect_compressibility(). If repetition ratio is low or input is already compact, passes through unchanged.
  2. Log & Structural Preprocessing: Performs count-preserving template deduplication ([A-N] exemplar), ANSI stripping, and JSON array flattening.
  3. Multi-Signal Scoring & Graph Closure: Computes lexical, entity, structural, and graph-expansion scores.
  4. Defensive Invariant Checks: Verifies that active errors, stack traces, and code fences are preserved intact (error-signal invariant).
  5. Lost-in-the-Middle Reordering (reorder_best): Places top-ranked evidence at the front and second-best at the tail to optimize LLM decoder attention.
  6. Self-Check Diagnostic Verifier: Performs post-compaction validation:

$$\text{Confidence Score} = 0.50 \cdot \text{Recall(ent)} + 0.30 \cdot \text{Recall(kw)} + 0.20 \cdot \text{CriticalLineRatio}$$


🧩 Supported Modalities & Preprocessing Engine

Modality Specialized Processing Typical Reduction
Terminal & Build Logs Count-preserving deduplication, ANSI stripping, stack frame collapse 90–96%
JSON Schemas & APIs Key-preserving array crushing, bounded nesting, exact selector matching 75–88%
NDJSON / JSONL Per-record validation and exact raw-line selection 80–99%
CSV / TSV Quote-aware record framing, header retention, embedded-newline safety 70–95%
Markdown Heading ancestry and atomic fenced-code sections 60–90%
YAML Parent-key plus matching-subtree retention; conservative decline on ambiguous prose 60–90%
XML / HTML Exact line-oriented child selection with preserved wrappers; unsafe multiline shapes decline 50–90%
SQL Quote/comment/dollar-string-aware statement selection 60–95%
INI / TOML-style sections Complete matching-section retention 60–95%
Git Dumps & Commit Logs Commit hash retention, author/subject line-record filtering 80–92%
Code Diffs & Patches Structural fence invariant, changed-line hunk isolation 65–80%
Multilingual / Mixed Direction Grapheme-safe logical-order matching across 20 script/language families 60–95%
Vertical OCR Columns Blank-column framing with logical-source-order matching 50–90%

🌍 Unicode, Direction & Vertical Text

Tameru never applies visual bidi reordering and never rewrites output into a different normalisation form. Original substrings remain the source of truth. A separate NFKC/case-folded matching shadow is used only for search.

  • Extended grapheme tailoring keeps combining marks, variation selectors, emoji modifiers, flags, Indic virama sequences and ZWJ/ZWNJ sequences atomic.
  • Direction profiles report ltr, rtl, mixed, or neutral using Unicode bidi classes. Explicit controls and overrides are counted separately.
  • Script-aware query units cover Arabic, Hebrew, Persian/Urdu, Devanagari, Bengali, Tamil, Telugu, Thai, Lao, Khmer, Myanmar, Han, Kana, Hangul, Greek, Cyrillic, Armenian, Georgian, Ethiopic and Mongolian families.
  • Space-free scripts use bounded grapheme n-grams; spaced scripts use logical word runs. ASCII keeps the original compiled-regex fast path.
  • CSS writing-mode and one/two-grapheme OCR columns are metadata hints only; they never reverse or rotate source text.

See docs/industrial-compaction-v1.2.md for the standards, tailoring decisions and invariants.


🏭 Industrial Limits & Format Contracts

Default limits are deliberately conservative and configurable per call:

Limit Default
Input characters 8,000,000
Logical lines 250,000
Structured records 100,000
Characters per record 1,000,000
Fields per structured record 4,096
Unicode profile sample 16,384 chars at each edge
Bidi controls 10,000
Bidi overrides 128

If a hard limit, malformed surrogate, parser invariant, or structure check fails, Tameru returns the caller's exact text before CCR writes or lossy work. Adapters only activate when they produce a shorter structurally valid exact subset; otherwise the v1.1 scorer remains the fallback.

from tameru import IndustrialLimits
from tameru.compress_context import compress_context

limits = IndustrialLimits(
    max_input_chars=4_000_000,
    max_records=50_000,
    max_bidi_overrides=16,
)
result = compress_context(context_data, query, limits=limits)
profile = result.receipt["industrial"]["profile"]

Equivalent CLI controls are available as --max-input-chars, --max-lines, --max-records, --max-record-chars, --max-fields, --max-profile-chars, --max-bidi-controls, and --max-bidi-overrides.


📊 Head-to-Head Competitive Benchmark

Tested across 17 standardized production-QA fixtures containing multi-hop reasoning, temporal overrides, noisy logs, and structured sample text:

Compaction System Gold Fact Recall Latency (avg) Cost / 1k Ops Deterministic Dependencies
⚡ Tameru (v1.3.0) 17 / 17 (100%) 4–956 ms observed; <1.2 s at 500 KB $0.00 100% Yes Python stdlib
BM25 / Vector RAG Baseline 12 / 17 (70.6%) ~400 ms $0.02–$0.05 No Vector DB + Embeddings
Abstractive LLM Summarizer 7 / 17 (41.2%) ~2,500 ms $1.50–$3.00 No (Stochastic) Auxiliary LLM API
Uncompressed Baseline 17 / 17 (100%) 0 ms Full Tokens Yes None

⚔️ Live comparison — six arms, one corpus (jev-1.13.0, Sep 2026)

Same 12-case QA corpus, measured via benchmarks/jev_comparison.py. Arms: real TypeSafe JEV via LCC (typed noul keep-probability questions — the real System One protocol), its local Laya decision model (laya-multilingual, CPU) and mechanical fallback, headroom-ai's Rust TextCrusher, and a stdlib TF-IDF retrieval baseline:

Metric Tameru (v1.3.0) lcc → JEV lcc → Laya lcc mechanical headroom TF-IDF
Gold retention 12/12 11/11 † 12/12 12/12 10/12 11/12
Forbidden distractors kept 0 5 5 5 4 5
Deterministic ✅ byte-identical ❌ ✅ ✅ ✅ ✅
Median latency 10.75 ms 591 ms † 8,675 ms 4.5 ms 0.5 ms <1 ms
Mean savings 81.0% 63.6% 10.7% 48.4% 49.6% 55.2%
Cost / runs local $0 / ✅ API-priced / ❌ $0 / ✅ $0 / ✅ $0 / ✅ $0 / ✅

† JEV skipped the 4,000-block perf case (API cost); Laya ran it — ~56 min on CPU.

The decisive gap isn't relevance — most judges kept the gold. It's admissibility: every non-Tameru arm kept planted distractors, including a stale superseded config and EXCLUDED-HOST inside a block labeled "UNTRUSTED SAMPLE". A keep-score — JEV probability, Laya decision, BM25, or TF-IDF — has no trust, supersession, or exclusion model; Tameru encodes all three. Full table, per-case detail, and methodology: benchmarks/COMPARISON.md. Ecosystem + paper digest: docs/RESEARCH.md.


🧪 The Production QA Battery (v3)

The tests/ directory contains an exhaustive production-QA suite:

  • test_arabic_and_sql_sink.py: Non-Latin script preservation and SQL tabular sink compaction.
  • test_adjacent_v010.py: Compaction audit logging, pinned sink regions, and versioned context receipts.
  • test_hardening_pass_v09.py: NTK template deduplication, progress-bar stripping, and stack frame collapse.
  • test_semantic_gates.py: Counterfactual distractor masking and graph closure validation.
  • test_production_qa_v3.py: Cross-seed determinism validation (PYTHONHASHSEED=0 vs PYTHONHASHSEED=1).

Run the full battery:

python -m unittest discover -s tests

Current v1.3.0 release verification:

  • 326 passed, 14 skipped with pytest.
  • 15/15 production-QA battery cases passed (benchmarks/run_battery.py), including Japanese, Chinese and Arabic needle cases with zero forbidden-distractor leaks.
  • benchmarks/threshold_sweep.py: every budget_ratio operating point from 0.1–0.7 keeps all gold with 0 leaks — a clean 77%→25% savings frontier.
  • Six-arm live comparison (Tameru / JEV / Laya / mechanical / headroom / TF-IDF): only arm with 12/12 gold + 0 leaks + deterministic output — table below.

🔄 Reversible Compaction & ARC Citations (CCR Store)

When Tameru drops content blocks, it emits content-addressed ARC Citations:

[A 7f8a9b] "2026-08-25 02:00:00 [INFO] Worker started"...[48 lines omitted]..."2026-08-25 02:05:00 [INFO] Complete"

If an agent subsequently determines that it needs the omitted content, it can query the local Content-Centric Retrieval (CCR) store:

from tameru.compress_context import retrieve

full_context = retrieve("7f8a9b2c4d6e8f0011223344")

CCR is opt-in per integration path. The standalone API keeps its reversible default, while the Hermes live tool-pruning adapter disables CCR because it has no retrieval path and tool payloads may contain secrets. Expired valid records are swept whenever a new CCR record is written; operators can also call sweep_ccr_cache() directly. CCR directories and records are created with owner-only permissions where the platform supports POSIX modes.


🔌 Hermes Context-Engine Plugin Integration

Tameru integrates into Hermes Agent as a pluggable context engine — no Hermes source changes:

  • Standalone plugin — 0xWhiteMage/hermes-tameru-plugin: copy into ~/.hermes/plugins/tameru/ or pip install it (auto-discovered via the hermes_agent.plugins entry point).
  • Activate in config.yaml:
    context:
      engine: tameru

🚀 Quickstart & Installation

Option 1: Standard Installation (Zero Dependencies)

pip install -e .

Option 2: With Opt-in Semantic Embeddings

pip install -e '.[semantic]'

Option 3: LLM / Agent Prompt Setup

To equip an autonomous agent with Tameru, paste this prompt into your agent's system prompt or session:

Please install and enable the Tameru compaction system for high-performance context management:
1. Clone or vendor tameru into your environment.
2. In your context configuration, select 'tameru' as the active context engine.
3. Use tameru.compress_context(context, query) before passing large tool outputs to your context window.

💻 Python API & Agent Integration

from tameru.compress_context import compress_context

context_data = (
    "2026-08-25 02:00:00 [INFO] System boot sequence started.\n"
    + ("2026-08-25 02:01:00 [INFO] Routine worker heartbeat.\n" * 500)
    + "2026-08-25 02:14:00 [ERROR] DB connection failed on host db-prod-01: Port 5432 unreachable.\n"
)

query = "Why did the database connection fail?"

result = compress_context(context_data, query)

print(f"Compressed Output:\n{result.compressed_text}\n")
print(f"Token Reduction: {result.tokens_saved_pct:.1f}%")
print(f"Fail-Open Triggered: {result.fail_open}")
print(f"Confidence Diagnostic: {result.verifier}")

strategy="summarise" remains fail-open and can be configured per call with summary_endpoint, summary_models, and summary_timeout. Equivalent environment variables are TAMERU_SUMMARY_ENDPOINT, TAMERU_SUMMARY_MODELS (comma-separated), and TAMERU_SUMMARY_TIMEOUT (seconds). The timeout is one total retry budget across all candidate models, not a fresh timeout per model.

Summary endpoints are restricted to loopback by default. Set TAMERU_SUMMARY_ALLOW_REMOTE=1 or pass summary_allow_remote=True only when the integration has explicitly approved sending context to that endpoint.

When using decision_cache, first-sight keep/drop outcomes are replayed on later turns. The cache preserves the oldest prefix and is bounded to 4,096 block decisions. Prefix stability intentionally outranks a later fixed-mode budget_ratio; if frozen keeps exceed that ratio, the engine preserves the frozen prefix and reports freeze cache capacity reached once the cache can no longer learn new blocks.

Additional caller controls, inspired by transcript-level pruners:

  • pin_recent=N unconditionally keeps the opening block and the newest N blocks — a working-context floor that no selector, freeze, supersession, or trust path may evict.
  • min_savings_ratio=R (default 0.10) fails open when achieved savings fall below R, so callers can decline a compaction that did not shrink enough; 0 disables the floor.
  • degraded_view=True rescues input size-limit breaches (max_input_chars, max_lines): instead of failing open, each oversized block is scored on a bounded head+tail view while emitted output stays byte-exact. Safety limits (bidi controls, malformed surrogates, oversize query) remain hard failures, and work stays bounded by a 4× grace factor over the configured limits. The receipt reports degraded_view: true and the risk floor becomes medium. CLI: --degraded-view.

Further ecosystem-research hardening (deterministic, always-on unless noted):

  • Secrets screen: if the input contains probable credentials (private key blocks, AKIA…, GitHub/Slack/sk- tokens, JWTs, long quoted password/api_key-style assignments), the CCR store is skipped — archival would persist secrets to disk at rest. The result reports ccr: None, emits no [CC-Retrieve:] marker, and the reason list notes ccr skipped: secret material detected.
  • Recursion guard: input already carrying Tameru output markers (<compressed_context, [CC-Retrieve:) returns unchanged with policy_name="local-noop-recursion" — compacting compacted output would nest wrappers and could drop the pointer that makes earlier drops recoverable.
  • Compressible-subset budgeting: in mode="fixed", pinned blocks are immovable and no longer consume the budget_ratio they outrank; the ratio governs only compressible tokens. With no pins, budgeting is unchanged.
  • CCR recall tooling: list_ccr(ccr_dir, offset=, limit=) lists live records newest-first (hash, stored_at, ttl, chars, preview), and retrieve(hash, offset=, limit=) paginates large originals.
  • Selection transparency: every receipt reports selection — which selector path decided the keep-set (needle, floor, floor-saturated, line-records, fixed, or a *-failopen variant), so silent degradation is visible to callers.
  • strategy="auto" runs the progressive ladder: extract first, then escalate to summarise only when extraction fails open (ambiguity, saturated floors, min_savings_ratio undershoot). summarise still falls back to extract on any LLM failure, so the worst case is one extra deterministic pass. CLI: --strategy auto.

Second research round (v1.3.0 — lcc internals, SelfCompact, PAACE, TPC, Compactor):

  • Dependency closure (sufficiency restore): after selection, a dropped block that shares a rare term with a kept block AND carries a qualifier cue (except, unless, only, however, until, if, …) or a definition cue (is defined as, refers to, namely, …) is restored — cutting it would invert the meaning of what survives (e.g. a kept "throughput is nominal" whose dropped sibling says "except during maintenance"). Capped at 8 restorations, never resurrects trust-risk or frozen-drop blocks, supersession can still evict stale restorations, and mode="fixed" restores only within the caller's hard budget. Receipts report sufficiency_restored: [block ids].
  • Qualifier-aware trim refusal: _crush_value no longer truncates a long JSON string whose cut tail carries a qualifier cue — a longer safe value beats a shorter misleading one.
  • Timing gate (transcript adapter): tameru.transcript.trajectory_gate suppresses the Tameru prune pass mid-derivation (pending tool calls) and on stuck loops (the last 3 assistant turns issued identical calls — diagnose, don't erase evidence). On by default via apply_extractive_tool_prune(..., timing_gate=True); it only ever suppresses that prune step — downstream host compaction is unaffected.
  • Plan-aware multi-query: compress_context(ctx, [q1, q2, ...]) scores blocks against the union of current + planned tasks.
  • Derived objective (derive_query=True, CLI --derive-query): an empty/generic query normally fails open; with derivation the engine infers a conservative task descriptor from the document's own recurring rare terms (deterministic frequency statistics, bounded sampling). Receipts mark query_source: "derived" + derived_terms, the risk floor becomes medium, and documents with no stable term structure still fail open.
  • Compression ceiling: inspect_compressibility() now reports guaranteed_savings_pct + ceiling_class (dedupe-heavy / moderate / sparse) — the share removable by pure dedupe before any relevance judgement, so callers can see how much headroom a given context actually has.
  • benchmarks/threshold_sweep.py: replays the QA corpus across budget_ratio values and reports the gold/leak/savings frontier — operating points chosen on evidence, not inherited.
  • Multilingual coverage: zh_needle + ar_needle battery cases — compression tuned on English silently fails other scripts.

Harness integration

Tameru ships harness-agnostic. Three integration paths, cheapest first — see harnesses/README.md for the full contract:

  • Python API — pip install tameru-compaction-system, then call compress_context(text, query) for raw payloads or tameru.transcript.apply_extractive_tool_prune(messages) for OpenAI-style role/content conversation lists (Hermes, OpenCode, Codex, and most agent loops share that shape).
  • CLI — tameru-compress reads a file or stdin, writes compacted text, and emits JSON stats with --stats; the universal shell-out for non-Python harnesses.
  • Vendored plugin — copy src/tameru/*.py into the plugin dir and keep it current with python scripts/sync_to_harness.py <plugin_dir> --manifest plugin.yaml. Upstream owns every module except __init__.py; the harness owns registration glue and the manifest.

📜 Changelog

All notable changes, version milestones, and migration notes are tracked in CHANGELOG.md.

Highlights:

  • v1.3.0: Typed-edge dependency restoration, transcript timing gate, plan-aware multi-query, qualifier-aware trim refusal, derive_query, compression-ceiling reporting, threshold-sweep tool, multilingual battery.
  • v1.2.1: JEV-inspired hardening — pin_recent, min_savings_ratio, degraded_view, secrets screen, recursion guard, compressible-subset budgeting, CCR recall tooling, selection receipts, strategy="auto".
  • v1.2.0: Bounded industrial preflight, logical-order Unicode across 20 script/language families, ten exact format adapters, deterministic receipt hashes and large-input SLOs.
  • v1.1.1: Comprehensive factual-retention, fail-open, cache-progression, public-metrics, and current-Hermes integration hardening.
  • v1.1.0: Production QA hardening, Docker progress recognition, CCR security & expiry sweep, bounded decision caching, loopback summary boundary.
  • v0.10.0: Compaction audit logs, pinned sink regions, versioned context receipts.
  • v0.9.0: NTK layer-1 log dedup, error invariants, progress-bar stripping, lost-in-the-middle reordering.
  • v0.8.0: Opt-in semantic tier with bi-encoders and cross-encoders.
  • v0.1.0 - v0.7.0: Core extractive engine, graph closure, temporal supersession, and tabular crushers.

🙏 Credits & Research Lineage

Tameru's design was inspired and sharpened by studying remarkable open-source projects, papers, and ideas:

Open-Source Projects

  • NTK / Neural Token Killer (Rust, MIT) — Template deduplication, error-signal invariant, progress-bar stripping, and stack frame collapse.
  • caveman-compression (MIT) — Executable test specification discipline and factual preservation validation.
  • leanctx (MIT) — Loss-tolerance routing and structural verbatim code invariants.
  • TwoTrim (Apache-2.0) — Lost-in-the-middle mitigation and attention edge anchoring.
  • clipforge-PAKT (MIT) — Pre-flight compressibility inspection.
  • fast-jev-compaction (MIT) — Positional pinning of edge turns, caller-side savings-worthiness gates, and judging on a degraded state while emitting exact spans.
  • dsh-argp — Compressible-subset retention budgeting: immovable atoms are excluded from the ratio's base.
  • dsh-jev-prune — Never degrade silently: report which selector path decided, and hard-exclude mutating actions by rule rather than by score.
  • dsh-jev-pre-compaction — Scan for secrets before archiving originals; lookahead pressure bands.
  • pi-lcm — Persistent append-only memory with self-serve grep/describe/expand recall tools.
  • lcc / Local Context Compiler (MIT) — Typed context-graph closure, post-drop sufficiency verification, qualifier-aware safe trimming, and a well-built JEV/local-model client used as our benchmark harness.
  • headroom — Closest production analog: extractive crushers + reversible CCR + cache alignment; its CJK token-pricing fix prompted our (passing) audit.
  • Waxmell114514/jev-compaction — Score-only reversible compaction and the threshold-replay tuning methodology behind benchmarks/threshold_sweep.py.
  • SelfCompact — The closed-unit/progress/not-stuck compaction-timing rubric, ported as tameru.transcript.trajectory_gate.
  • Claude Code compaction teardown (anneheartrecord/claude-code-docs) — Progressive micro→session→full tiers, recursion guards on compaction agents, and circuit breakers after repeated failures.

Research Lineage

  • LongLLMLingua / LLMLingua-2 (Microsoft, ACL 2024) — Question-aware compression and the preserve/discard token-classification formulation.
  • Lost in the Middle (Liu et al., 2023) — Positional attention decay in decoder transformers.
  • NoLiMa (2025) — Long-context lexical distractor blind spots.
  • Context Rot (Chroma, 2025) — The empirical case for deliberate compression: models degrade as input grows even on simple tasks.
  • Provence (2025) — "Zero sentences kept" as a legal answer; pruning and ranking as one operation.
  • PAACE (2025) — Plan-aware selection for next-k tasks, ported as multi-query union scoring.
  • TPC (2025) — Task-descriptor objectives for question-free compression, ported deterministically as derive_query.
  • Compactor (2025) — Context-calibrated compression ceilings, ported as inspect_compressibility guaranteed-savings reporting.
  • Lost in Compression (2026 audit) — Cross-lingual failure modes of English-tuned compressors; drove the multilingual battery.
  • ARC Citations — Reversible content-addressed reference anchors.

Full source-by-source digest: docs/RESEARCH.md.


🤝 Community & Support


📄 License

Released under the MIT License.

About

Deterministic, query-aware context compaction — harness-agnostic, local, reversible, fail-open, multilingual, and zero-LLM in the extractive path.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages