The user's core requirement is "the AI builds from start to finish without breaking." The dominant thing that breaks an end-to-end build is not a weak model — it is context rot: as the input grows, the model uses it less reliably and loses the thread. This document is the package's answer to that failure mode. Continuity is engineered, not assumed.
Chroma's Context Rot report (Hong, Troynikov, Huber, 2025) evaluated 18 SOTA models (GPT-4.1, Claude 4, Gemini 2.5, Qwen3) and concluded: "models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows." RULER (arXiv:2404.06654) and NoLiMa (arXiv:2502.05167) corroborate from independent angles. The practical implication, in Anthropic's words: "ASSUME INTERRUPTION: your context window might be reset at any moment, so you risk losing any progress that is not recorded in your memory directory."
The mitigation is not a bigger window. It is progressive disclosure plus external memory. The spec package is structured to deliver both.
The taxonomy splits the spec into small files precisely so the
builder loads constitution.md + AGENTS.md up front and then glob/greps
for the one REQ-007 it needs. It never holds the whole spec. The
cross-link topology makes neighbor retrieval a
single hop. This is the primary defense: most of the spec stays out of the
window until needed.
The agentic memory file is append-only, date-stamped, and lives on disk, not in the window. On resume, the builder reads its tail to reconstruct state. Mandatory sections (template: changelog.md):
- Current status — which
T-NNNis in progress, which gate is next. - Completed — tasks done this run, with commit refs.
- Failed approaches & why — the dead-end ledger. This is the single most valuable section. It stops a resumed agent from re-trying what already failed.
- Open questions — unresolved
[NEEDS CLARIFICATION]items. - Accuracy / metric snapshots — at checkpoints, so regressions are visible.
tasks.md carries a status per task; tasks.json mirrors it as
machine-readable state. A reset agent reads the DAG, finds the first pending
task whose depends_on are all done, and resumes there. No re-derivation of
"where was I" from conversation history.
Anthropic notes models are less likely to re-write JSON than Markdown. So
append-only, high-stakes state — the task DAG, gate results, the
traceability matrix — is mirrored to JSON (tasks.json, traceability.json,
gate-results.json). Prose the agent should freely edit stays Markdown. The
frontmatter
bridges the two.
Chroma's data shows degradation well before the hard context limit. So
planning sessions should compact at ~70–80% of the window (more aggressive
than the typical ~95% default), writing a summary to NOTES.md first so nothing
is lost. The order is fixed: persist to memory file → compact → reload
constitution + memory tail.
When a builder agent starts (or restarts mid-build), it runs this exact sequence before touching code:
1. Load constitution.md + AGENTS.md (Layer 0, always)
2. Read CHANGELOG.md tail (status + failed approaches)
3. Read tasks.json (find next ready task)
4. For the chosen T-NNN, load only its edges:
spec.md#req-* · design.md#cmp-* · data-model.md#entity-* ·
api-contract.md#op-* · active overlay sections
5. Implement → test → tick AC → append CHANGELOG.md
6. If ambiguity: write [NEEDS CLARIFICATION], route to clarify loop, do not guess
This protocol is what makes the package resumable. It is also the basis for the
continuity axis of the evaluation: force a context reset
every N tokens and measure whether CHANGELOG.md + tasks.json are sufficient
for the agent to finish with ≤10% extra tokens.
| Without continuity engineering | With it |
|---|---|
| Re-asks the user what was decided last week | Reads ADR + CHANGELOG.md |
| Re-tries a refactor that already failed | Reads "failed approaches & why" |
| Loses track of which tasks are done | Reads tasks.json status |
| Holds the whole 60-page PRD, degrades | Holds one node + edges |
| Compacts too late, drops a decision | Persists to memory, compacts at 70–80% |
Chroma Context Rot: How Increasing Input Tokens Impacts LLM Performance (2025); Anthropic Effective context engineering for AI agents (2025), Effective harnesses for long-running agents (2025), Claude memory-tool documentation; RULER (arXiv:2404.06654); NoLiMa (arXiv:2502.05167). Note: context-rot reports are vendor/engineering-published, corroborated by peer-reviewed RULER and NoLiMa; specific token thresholds should be re-measured per model the runtime ships against.