Skip to content

Latest commit

 

History

History
110 lines (86 loc) · 5.36 KB

File metadata and controls

110 lines (86 loc) · 5.36 KB

Continuity Engineering: Surviving the Context Reset

The user's core requirement is "the AI builds from start to finish without breaking." The dominant thing that breaks an end-to-end build is not a weak model — it is context rot: as the input grows, the model uses it less reliably and loses the thread. This document is the package's answer to that failure mode. Continuity is engineered, not assumed.

The failure mode, measured

Chroma's Context Rot report (Hong, Troynikov, Huber, 2025) evaluated 18 SOTA models (GPT-4.1, Claude 4, Gemini 2.5, Qwen3) and concluded: "models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows." RULER (arXiv:2404.06654) and NoLiMa (arXiv:2502.05167) corroborate from independent angles. The practical implication, in Anthropic's words: "ASSUME INTERRUPTION: your context window might be reset at any moment, so you risk losing any progress that is not recorded in your memory directory."

The mitigation is not a bigger window. It is progressive disclosure plus external memory. The spec package is structured to deliver both.

Five continuity mechanisms the package provides

1. Progressive disclosure (load small, retrieve just-in-time)

The taxonomy splits the spec into small files precisely so the builder loads constitution.md + AGENTS.md up front and then glob/greps for the one REQ-007 it needs. It never holds the whole spec. The cross-link topology makes neighbor retrieval a single hop. This is the primary defense: most of the spec stays out of the window until needed.

2. External memory file (CHANGELOG.md / NOTES.md)

The agentic memory file is append-only, date-stamped, and lives on disk, not in the window. On resume, the builder reads its tail to reconstruct state. Mandatory sections (template: changelog.md):

  • Current status — which T-NNN is in progress, which gate is next.
  • Completed — tasks done this run, with commit refs.
  • Failed approaches & why — the dead-end ledger. This is the single most valuable section. It stops a resumed agent from re-trying what already failed.
  • Open questions — unresolved [NEEDS CLARIFICATION] items.
  • Accuracy / metric snapshots — at checkpoints, so regressions are visible.

3. Checkpointed task state in the DAG

tasks.md carries a status per task; tasks.json mirrors it as machine-readable state. A reset agent reads the DAG, finds the first pending task whose depends_on are all done, and resumes there. No re-derivation of "where was I" from conversation history.

4. JSON for state the model must not corrupt {#json-for-state}

Anthropic notes models are less likely to re-write JSON than Markdown. So append-only, high-stakes state — the task DAG, gate results, the traceability matrix — is mirrored to JSON (tasks.json, traceability.json, gate-results.json). Prose the agent should freely edit stays Markdown. The frontmatter bridges the two.

5. Compaction before the cliff

Chroma's data shows degradation well before the hard context limit. So planning sessions should compact at ~70–80% of the window (more aggressive than the typical ~95% default), writing a summary to NOTES.md first so nothing is lost. The order is fixed: persist to memory file → compact → reload constitution + memory tail.

The resume protocol

When a builder agent starts (or restarts mid-build), it runs this exact sequence before touching code:

1. Load constitution.md + AGENTS.md            (Layer 0, always)
2. Read CHANGELOG.md tail                       (status + failed approaches)
3. Read tasks.json                              (find next ready task)
4. For the chosen T-NNN, load only its edges:
     spec.md#req-*  ·  design.md#cmp-*  ·  data-model.md#entity-*  ·
     api-contract.md#op-*  ·  active overlay sections
5. Implement → test → tick AC → append CHANGELOG.md
6. If ambiguity: write [NEEDS CLARIFICATION], route to clarify loop, do not guess

This protocol is what makes the package resumable. It is also the basis for the continuity axis of the evaluation: force a context reset every N tokens and measure whether CHANGELOG.md + tasks.json are sufficient for the agent to finish with ≤10% extra tokens.

Anti-patterns this prevents

Without continuity engineering With it
Re-asks the user what was decided last week Reads ADR + CHANGELOG.md
Re-tries a refactor that already failed Reads "failed approaches & why"
Loses track of which tasks are done Reads tasks.json status
Holds the whole 60-page PRD, degrades Holds one node + edges
Compacts too late, drops a decision Persists to memory, compacts at 70–80%

Sources

Chroma Context Rot: How Increasing Input Tokens Impacts LLM Performance (2025); Anthropic Effective context engineering for AI agents (2025), Effective harnesses for long-running agents (2025), Claude memory-tool documentation; RULER (arXiv:2404.06654); NoLiMa (arXiv:2502.05167). Note: context-rot reports are vendor/engineering-published, corroborated by peer-reviewed RULER and NoLiMa; specific token thresholds should be re-measured per model the runtime ships against.