Skip to content
Merged
Show file tree
Hide file tree
Changes from 4 commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS_USAGE_INSTRUCTIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

Canonical usage instructions for AI agents using FAVA Trails MCP tools. Other docs reference this file — keep it up to date.

> **Auto-injected:** Core guidance from this file is automatically injected via the MCP server's `instructions` field at session init — no manual setup required. The full version below is also available on-demand via the `get_usage_guide` tool. This file is the canonical source for both.
> **Session-init subset:** On the default `full` MCP surface, the server `instructions` field is a maintained subset of this file, not a verbatim inject. Session-start recall examples in that subset must match the fenced examples below. `FAVA_TRAILS_MCP_SURFACE=compact` sends a short pointer instead. This file is the canonical source and is returned verbatim by `get_usage_guide`. Compact does not make instructions a shared memory store.

## Governed recall

Expand Down
1 change: 1 addition & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@ All notable changes to FAVA Trails are documented here.
## Unreleased

### Added
- `FAVA_TRAILS_MCP_SURFACE=compact` advertises shorter initialize instructions and tool descriptions and omits list-time `outputSchema`, with `get_usage_guide` as on-demand protocol. Default remains `full`. `fava-trails measure-mcp-context` records tokenizer-labeled session-init size, loads a frozen issue #104 tested-release artifact (not relabeled from the current SDK), and the same-task comparison runs through `mcp.Client` sessions. Prompt-coverage gaps are instruction scans, not observed client skips; missing-scope recovery selects an exact returned path and retries recall. Full initialize instructions are a maintained subset of `AGENTS_USAGE_INSTRUCTIONS.md`, not a verbatim inject. See [docs/mcp-context-overhead.md](docs/mcp-context-overhead.md). Addresses #104.
- `fava-trails register` prints native MCP registration using an ordinary `FAVA_TRAILS_AGENT_ID`, the resolved executable, and the intended data repository. Unresolved executables fail unless `--executable` names an existing executable file. `--write` is an explicit client-config opt-in (atomic write, `.bak` backup, files opened with the final mode before content is written, modes capped at `0600` while stricter existing modes are kept, new files `0600`, non-writable existing configs refused). `--verify` labels a direct MCP smoke test, client-config inspection, and MCP Inspector config-load (`inspector_config_load`). It does not claim Claude Code/Desktop loaded the registration. Failures stay distinct (`inspector_unavailable`, `inspector_invocation_failed`, `config_load_failed`, `server_spawn_failed`, `server_initialize_failed`, `inspector_failed`, stale runtime path, registration not loaded). Native-session evidence that a client loaded Claude-shaped `mcpServers` config is `test_native_client_registration_loads_and_initializes`.
- Bounded obvious-secret preflight before save, update, supersede, and promotion persist or transmit. Supported high-confidence patterns are refused with a safe explanation, including nested caller-controlled metadata and relationships after hook mutation, and the complete MCP request (tool name plus arguments) before schema validation, logging, lookup, auto-initialization, or JJ operations. Assembled Trust Gate result metadata is scanned before governance persist. Nested walks deeper than 32 fail closed. Block logs use a fixed message without pattern ids. Legacy matching drafts are left unchanged and are not sent for review. Documents data flow and detection limits; does not claim complete DLP. Fixes #102.
- **Trust Gate data-egress disclosure (issue #101):** `describe_trust_gate_egress`
Expand Down
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -388,6 +388,7 @@ Environment variables:
| `FAVA_TRAILS_DATA_REPO` | Server | Root directory for trail data (monorepo root) | `~/.fava-trails` |
| `FAVA_TRAILS_DIR` | Server | Override trails directory location (absolute path) | `$FAVA_TRAILS_DATA_REPO/trails` |
| `FAVA_TRAILS_SCOPE_HINT` | Server | Broad scope hint baked into tool descriptions | *(none)* |
| `FAVA_TRAILS_MCP_SURFACE` | Server | `full` (default) or `compact` advertised instructions/tool text | `full` |
| `FAVA_TRAILS_SCOPE` | Agent | Optional process override for project scope. Read if set; `fava-trails init` does not write application `.env` files unless `--write-env` is passed. | *(none)* |
| `OPENROUTER_API_KEY` | Server | Default Trust Gate API key env (OpenRouter). Override the env var *name* via `trust_gate_api_key_env` / legacy `openrouter_api_key_env` in `config.yaml`. | *(none — required for `propose_truth` when using llm-oneshot)* |

Expand Down Expand Up @@ -527,6 +528,7 @@ uv run pytest --cov # with coverage
- [AGENTS_SETUP_INSTRUCTIONS.md](AGENTS_SETUP_INSTRUCTIONS.md) — Data repo setup, config reference, trust gate prompts, lifecycle hooks
- [protocols/secom/README.md](src/fava_trails/protocols/secom/README.md) — SECOM compression protocol: config, models, WORM architecture
- [docs/fava_trails_faq.md](docs/fava_trails_faq.md) — Detailed FAQ for framework authors and ML engineers
- [docs/mcp-context-overhead.md](docs/mcp-context-overhead.md) — Measured MCP session-init overhead, compact surface, enforcement vs prompt
- [docs/secret-preflight.md](docs/secret-preflight.md) — Bounded credential preflight: data flow, detection limits, false positives

## Contributing
Expand Down
157 changes: 157 additions & 0 deletions docs/mcp-context-overhead.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,157 @@
# MCP context overhead

This note records how FAVA Trails measures advertised MCP session-init text, what a
compact surface changes, and which learning-workflow steps the server actually
enforces. Figures below are for one serialization and one tokenizer. They are not
a universal client token cost.

## How to measure

```bash
fava-trails measure-mcp-context --surface both
```

The command serializes:

- initialize `instructions`
- `tools/list` items as this server advertises them (JSON, compact separators)

It records FAVA package version and git commit of the measured checkout
(candidate), a frozen issue #104 tested-release artifact (`6c5278a40a86246014901a88417f3455a46cdfcc`,
not re-measured or relabeled from the current SDK), tokenizer name, MCP Python SDK
`mcp.Client` version, enabled tool names, `lazy_loading` (always `false`: every
tool is listed at `tools/list`), and whether the cost recurs. Instructions are
sent once per `initialize`. `tools/list` is sent once per list request; typical
clients list once per session and only re-pay the cost if they refresh the catalog.

Default tokenizer: `chars/4 heuristic` (`ceil(character_count / 4)`). If
`tiktoken` is installed, `cl100k_base` is recorded as an optional extra. Neither
figure is a client invoice.

## Provenance

| Checkout | Role | FAVA version | Git commit | Client |
| --- | --- | --- | --- | --- |
| Issue #104 source review baseline | tested release (frozen artifact) | 0.6.1 | `6c5278a40a86246014901a88417f3455a46cdfcc` | `mcp.Client` 2.2.0 (frozen) |
| This branch | candidate | 0.6.1 | current `git rev-parse HEAD` | live `mcp.Client` (currently 2.2.0) |

The tested release had no compact surface. Its full-surface payload, enabled tool
names, lazy-loading flag, recurrence, tokenizer, and client version are stored in
`src/fava_trails/issue_104_tested_release.json`. `fava-trails measure-mcp-context`
loads that file and does not overwrite client/SDK fields from the current
environment. Reproduce by checking out `6c5278a` and serializing advertised
initialize instructions plus `tools/list` JSON with the chars/4 heuristic.

## Recorded baseline

Tokenizer `chars/4 heuristic`, all 17 tools enabled, no lazy loading, no `tiktoken`.

Tested release (`6c5278a`, full surface only):

| Instructions tokens | tools/list tokens | Session-init tokens | Session-init chars | `get_usage_guide` chars / tokens |
| ---: | ---: | ---: | ---: | ---: |
| 961 | 5451 | 6412 | 25647 | 10077 / 2520 |

Candidate (this head, live `mcp.Client`):

| Surface | Instructions tokens | tools/list tokens | Session-init tokens | Session-init chars |
| --- | ---: | ---: | ---: | ---: |
| full (default) | 1154 | 5793 | 6947 | 27782 |
| compact | 187 | 3179 | 3366 | 13458 |

`get_usage_guide` body on this candidate (on demand, not in session-init): 2884
heuristic tokens (11536 chars). An evaluator previously estimated about 6000 tokens
of schemas and instructions versus about 1600 for a committed agent guide; that
estimate was client-specific and is not reproduced here as a universal number.

Budget, from the candidate full session-init baseline: compact session-init tokens
must be ≤ 70% of full under the same tokenizer. This run: 3366 / 6947 ≈ 0.48. Met.

Largest full-surface source is advertised `tools/list` JSON (schemas, then
descriptions), then initialize instructions. Compact therefore:

1. Shortens initialize instructions and points at `get_usage_guide`.
2. Shortens tool descriptions (drops duplicated session/promotion prose).
3. Omits advertised `outputSchema` on `tools/list`. Server-side validation still
uses `TOOL_DEFINITIONS`.

Input schemas, tool names, and authorization are unchanged.

## Compact surface

Default remains `full` (backward compatible). Opt in per MCP process:

```json
{
"mcpServers": {
"fava-trails": {
"command": "fava-trails-server",
"env": {
"FAVA_TRAILS_MCP_SURFACE": "compact"
}
}
}
}
```

Unknown values log as `full` at server start so a typo does not fail the process.
`fava-trails measure-mcp-context` rejects unknown surfaces.

## recall / save / promote comparison

Executed on both surfaces through in-process `mcp.Client` sessions (`mode="legacy"`
initialize handshake) against dedicated `Server` instances. The harness instantiates
the client; it does not call `handle_call_tool` directly. Task: observe initialize
instructions, follow only those instructions on a naive pass (no `get_usage_guide`),
then script invalid save, retry with content, missing-scope recall, `list_scopes`,
select an exact returned path, retry recall, authoring `recall`, and `propose_truth`
(Trust Gate review mocked). Scripted executed steps are recorded separately from
deterministic prompt-coverage scans of initialize text. Results:

| Check | full | compact |
| --- | --- | --- |
| Token usage (session-init, this tokenizer) | 6947 | 3366 |
| Discoverability of recall, save_thought, propose_truth, get_usage_guide, list_scopes | yes | yes |
| All 17 tools advertised | yes | yes |
| Input schemas | full | same |
| Advertised outputSchema | yes | omitted |
| Executed save_thought | ok | ok |
| Executed authoring recall after save | count 1 | count 1 |
| Executed propose_truth | ok | ok |
| Error recovery: missing scope | status error, then list_scopes + retry recall on returned path | same |
| Error recovery: save without content | failed, then retry ok | failed, then retry ok |
| Session-start recall trio in initialize text | yes | no; in `get_usage_guide` |
| `propose_truth` requested in initialize text | yes (mandatory wording) | yes (core loop; no “mandatory”) |
| Prompt-coverage gap (instruction scan, not a client choice) | none for session-start recall | session-start recall only |

Prompt-coverage indicators are a deterministic scan of initialize text, not
behavior observed from a client. Compact initialize still requests `propose_truth`
in the core loop; missing the word “mandatory” is not a skip. Compact omits the
session-start recall trio, which remains a coverage gap unless the client calls
`get_usage_guide` or injects its own guide. The server does not invoke those
steps on either surface. Missing-scope recovery selects an exact `list_scopes`
path and retries recall; a successful empty `list_scopes` is discovery attempted,
not recovery.

## What the server enforces vs prompt/client behavior

Server-enforced (same on both surfaces):

- `FAVA_TRAILS_AGENT_ID` identity match; caller `agent_id` cannot impersonate
- governed / authoring / history visibility
- operator-only tools (`diff`, `conflicts`, `rollback`, `forget`, `learn_preference`)
- writes require a configured agent identity
- input and output validation against `TOOL_DEFINITIONS`
- unpromoted drafts are not governed current records

Prompt/client behavior (not enforced by listing or instructions):

- calling `get_usage_guide`
- session-start recall of status/decisions/gotchas
- deciding work is “finalized” and calling `propose_truth`
- creating `.fava-trails.yaml` (do not write application `.env` files)
- whether the client shows initialize instructions or re-lists tools

Instructions do not provide reliable cross-session sharing. Sharing requires
`propose_truth` plus durable approval. Asking an agent to remember something in
the MCP instructions field does not make it available to the next session.
1 change: 1 addition & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -68,6 +68,7 @@ include = [

[tool.hatch.build.targets.wheel.force-include]
"AGENTS_USAGE_INSTRUCTIONS.md" = "fava_trails/AGENTS_USAGE_INSTRUCTIONS.md"
"src/fava_trails/issue_104_tested_release.json" = "fava_trails/issue_104_tested_release.json"

[tool.pytest.ini_options]
testpaths = ["tests"]
Expand Down
25 changes: 25 additions & 0 deletions src/fava_trails/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -2027,9 +2027,34 @@ def build_parser() -> argparse.ArgumentParser:
actions.add_argument("--rollback", action="store_true", help="Restore exact before images only if post-migration state still matches")
p_duplicates.set_defaults(func=cmd_duplicates)

p_measure = subparsers.add_parser(
"measure-mcp-context",
help="Measure serialized MCP instructions and tools/list for full and compact surfaces",
)
p_measure.add_argument(
"--surface",
choices=("full", "compact", "both"),
default="both",
help="Which advertised surface to measure (default: both)",
)
p_measure.set_defaults(func=cmd_measure_mcp_context)

return parser


def cmd_measure_mcp_context(args: argparse.Namespace) -> int:
"""Print a tokenizer-labeled measurement of MCP session-init payload size."""
import json

from .mcp_context import compare_surfaces, measure_mcp_context

if args.surface == "both":
print(json.dumps(compare_surfaces(), indent=2))
return 0
print(json.dumps(measure_mcp_context(args.surface), indent=2))
return 0


def cmd_duplicates(args: argparse.Namespace) -> int:
"""Local operator maintenance; dry-run is the default, never an MCP mutation."""
import asyncio
Expand Down
67 changes: 67 additions & 0 deletions src/fava_trails/issue_104_tested_release.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
{
"frozen": true,
"relabeling_forbidden": true,
"role": "release",
"package": "fava-trails",
"version": "0.6.1",
"git_commit": "6c5278a40a86246014901a88417f3455a46cdfcc",
"surface": "full",
"compact_surface_existed": false,
"client": {
"name": "mcp.Client",
"package": "mcp",
"version": "2.2.0",
"frozen": true,
"note": "Client identity frozen with this artifact. Do not overwrite with a later environment's MCP SDK version."
},
"tokenizer": {
"name": "chars/4 heuristic",
"not_universal": true
},
"lazy_loading": false,
"enabled_tools": [
"start_thought",
"save_thought",
"get_thought",
"propose_truth",
"recall",
"forget",
"sync",
"conflicts",
"rollback",
"diff",
"list_scopes",
"list_trails",
"learn_preference",
"update_thought",
"supersede",
"get_usage_guide",
"change_scope"
],
"tool_count": 17,
"recurrence": {
"instructions": "once per initialize",
"tools_list": "once per tools/list; typical clients list once per session and may re-list if they refresh the catalog. Cost recurs only when the client re-lists."
},
"instructions": {
"chars": 3843,
"utf8_bytes": 3867,
"tokens": 961
},
"tools_list": {
"chars": 21804,
"utf8_bytes": 21814,
"tokens": 5451
},
"session_init": {
"chars": 25647,
"utf8_bytes": 25681,
"tokens": 6412
},
"usage_guide_on_demand": {
"chars": 10077,
"utf8_bytes": 10123,
"tokens": 2520
},
"measured_how": "Frozen artifact: checkout 6c5278a40a86246014901a88417f3455a46cdfcc, serialize advertised initialize instructions plus tools/list JSON with compact separators, tokenize with chars/4 heuristic. Compact surface did not exist on that commit. Command loads this file and does not re-measure or relabel client/SDK fields."
}
Loading
Loading