Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 3 additions & 2 deletions AGENTS_SETUP_INSTRUCTIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -35,8 +35,9 @@ publication note.

The Trust Gate reviews thoughts before promotion using an LLM **or** an explicit
operator approval path. Provider selection is a **data-egress choice**: under the
shipped `llm-oneshot` policy, candidate thought content (plus selected redacted
metadata) is transmitted to the configured destination **before** a verdict
shipped `llm-oneshot` policy, candidate thought content (plus the selected
metadata fields; `agent_id` and `metadata.extra` are excluded) is transmitted to
the configured destination **before** a verdict
exists. A remote reject still means the content already left this process. There
is no automatic pass-through/off mode and no silent fallback to another provider.

Expand Down
3 changes: 3 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,9 @@ All notable changes to FAVA Trails are documented here.

## Unreleased

### Fixed
- Trust Gate data-egress disclosure no longer describes the transmitted metadata as "redacted": doctor, the MCP startup notice, and `trust_gate_egress` now state that the full scope-resolved prompt, the full candidate body, and the complete selected metadata fields are sent, with `agent_id` and `metadata.extra` excluded. Docs updated to match; regression coverage added for both `llm-oneshot` and `decisions` notices. Addresses issue #125.

### Added
- Explicit `decisions` Trust Gate policy: review and promote thoughts through OpenRouter's Decisions API (TypeSafe Jev) with one typed Noul question and an operator-configured threshold under `trust_gate_decisions_config` (`trust_gate_noul_question`, `trust_gate_noul_threshold`). Approves at-or-above threshold and rejects below through `propose_truth`, persists policy/provider/model/Noul-probability/threshold/timestamp/approval-kind provenance without fabricated reasoning, fails closed on invalid responses, auth/connection failures, and timeouts, and discloses secret-free destination/egress diagnostics in doctor and tunnel gateway startup. `llm-oneshot` remains the default. Implements #124.

Expand Down
4 changes: 2 additions & 2 deletions docs/fava_trails_faq.md
Original file line number Diff line number Diff line change
Expand Up @@ -88,7 +88,7 @@ Three critical differences.

### What is the Trust Gate, concretely?

The Trust Gate is a configurable validation step that can run before a draft is promoted into shared records. In the shipped default, it is a **synchronous** LLM (or explicit human) rubric review of the **current proposed record only**: the configured prompt plus that record's content and selected redacted metadata. It does **not** load the shared corpus, does **not** independently verify project facts, and is **not** a guarantee that hallucinations never enter shared truth.
The Trust Gate is a configurable validation step that can run before a draft is promoted into shared records. In the shipped default, it is a **synchronous** LLM (or explicit human) rubric review of the **current proposed record only**: the configured prompt plus that record's content and the selected metadata fields (`agent_id` and `metadata.extra` are excluded). It does **not** load the shared corpus, does **not** independently verify project facts, and is **not** a guarantee that hallucinations never enter shared truth.

The key design properties today:

Expand Down Expand Up @@ -211,7 +211,7 @@ Your actual data — the memory graph, the versioned repository, every thought y

What holds today: the engine does not embed your corpus in its source, does not collect product telemetry, and does not require a FAVA-hosted cloud. **Durable product state** lives only in Fuel under your controls; restarting the engine process does not move that corpus. Treat process-local caches (managers/hooks) as runtime convenience, not as a second source of truth.

What does **not** hold by default: **Trust Gate can egress content.** With the shipped OpenRouter default under `trust_gate: llm-oneshot`, `propose_truth` sends the proposed record content and selected redacted metadata to that provider **before** a verdict exists (rejection by a remote gate still happens after transmission). Run `fava-trails doctor` (or inspect `trust_gate_egress` on the tool response) for the effective destination/model and a plain summary of which fields are sent — API keys are never shown. To avoid remote egress today: point Trust Gate at a local OpenAI-compatible endpoint (fail-closed; no silent cloud fallback), **or** promote on an operator endpoint with `propose_truth(..., approval="human")` (requires `FAVA_TRAILS_OPERATOR=1` and a configured agent identity; no LLM transmission). There is no shipped config flag that disables LLM review globally; `trust_gate: human` is not implemented. Your corporate IP stays on your infrastructure only to the extent your Trust Gate provider, approval path, and hosting choices keep it there.
What does **not** hold by default: **Trust Gate can egress content.** With the shipped OpenRouter default under `trust_gate: llm-oneshot`, `propose_truth` sends the proposed record content and the selected metadata fields (`agent_id` and `metadata.extra` are excluded) to that provider **before** a verdict exists (rejection by a remote gate still happens after transmission). Run `fava-trails doctor` (or inspect `trust_gate_egress` on the tool response) for the effective destination/model and a plain summary of which fields are sent — API keys are never shown. To avoid remote egress today: point Trust Gate at a local OpenAI-compatible endpoint (fail-closed; no silent cloud fallback), **or** promote on an operator endpoint with `propose_truth(..., approval="human")` (requires `FAVA_TRAILS_OPERATOR=1` and a configured agent identity; no LLM transmission). There is no shipped config flag that disables LLM review globally; `trust_gate: human` is not implemented. Your corporate IP stays on your infrastructure only to the extent your Trust Gate provider, approval path, and hosting choices keep it there.

This separation also means you can update the engine independently of your data. Upgrading FAVA Trails's MCP server does not touch, migrate, or expose your repository.

Expand Down
2 changes: 1 addition & 1 deletion docs/fava_trails_onepager.html
Original file line number Diff line number Diff line change
Expand Up @@ -990,7 +990,7 @@ <h2>MCP-native. Seventeen tools.</h2>
<section class="reveal">
<div class="section-label">Enterprise Objections, Answered</div>
<h2>Your memory graph stays under your control</h2>
<p>FAVA Trails follows an <strong style="color: var(--accent);">Engine vs. Fuel</strong> architecture. The MCP server (<code style="font-family: var(--mono); font-size: 0.85em; color: var(--green);">fava-trails</code>) is an open-source engine process: it retains in-memory managers, backend handles, locks, loaded hooks/prompts, and config for the life of the process, but the <em>durable corpus</em> lives only in your separate Fuel directory (local-first or privately hosted). The engine does not embed your corpus in its source, collect product telemetry, or require a FAVA cloud. <strong>Default Trust Gate is different:</strong> under shipped <code style="font-family: var(--mono); font-size: 0.85em;">trust_gate: llm-oneshot</code> with the OpenRouter default, <code style="font-family: var(--mono); font-size: 0.85em;">propose_truth</code> sends the proposed record content and selected redacted metadata to that LLM provider. Avoid that path today by pointing Trust Gate at a local OpenAI-compatible endpoint, or by promoting on an operator endpoint with <code style="font-family: var(--mono); font-size: 0.85em;">propose_truth(..., approval="human")</code>. There is no shipped global “disable LLM review” config flag (<code style="font-family: var(--mono); font-size: 0.85em;">trust_gate: human</code> is unimplemented) — see issue #101 for further local-only hardening.</p>
<p>FAVA Trails follows an <strong style="color: var(--accent);">Engine vs. Fuel</strong> architecture. The MCP server (<code style="font-family: var(--mono); font-size: 0.85em; color: var(--green);">fava-trails</code>) is an open-source engine process: it retains in-memory managers, backend handles, locks, loaded hooks/prompts, and config for the life of the process, but the <em>durable corpus</em> lives only in your separate Fuel directory (local-first or privately hosted). The engine does not embed your corpus in its source, collect product telemetry, or require a FAVA cloud. <strong>Default Trust Gate is different:</strong> under shipped <code style="font-family: var(--mono); font-size: 0.85em;">trust_gate: llm-oneshot</code> with the OpenRouter default, <code style="font-family: var(--mono); font-size: 0.85em;">propose_truth</code> sends the proposed record content and the selected metadata fields (<code style="font-family: var(--mono); font-size: 0.85em;">agent_id</code> and <code style="font-family: var(--mono); font-size: 0.85em;">metadata.extra</code> are excluded) to that LLM provider. Avoid that path today by pointing Trust Gate at a local OpenAI-compatible endpoint, or by promoting on an operator endpoint with <code style="font-family: var(--mono); font-size: 0.85em;">propose_truth(..., approval="human")</code>. There is no shipped global “disable LLM review” config flag (<code style="font-family: var(--mono); font-size: 0.85em;">trust_gate: human</code> is unimplemented) — see issue #101 for further local-only hardening.</p>

<div class="ip-bullet">
<span class="icon">🔒</span>
Expand Down
2 changes: 1 addition & 1 deletion docs/secret-preflight.md
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ records.
| Preflight reject on save / update / supersede / MCP arguments | No new thought file; no new JJ/Git snapshot of the candidate; no canary-bearing trail directory | No | Scans the tool name plus argument object before schema validation or log/lookup, then body plus nested caller-controlled strings after hooks; safe error names the pattern id only; block logs use a fixed message |
| `save_thought` success | Draft markdown under `thoughts/drafts/` plus a JJ commit | Only if `push_strategy=immediate` later publishes the repo | This is persistence, not review |
| `update_thought` / `supersede` success | In-place rewrite or a new draft successor in history | Same as save | Supersede `reason` and copied metadata/relationships are also scanned |
| `propose_truth` LLM review | Already on disk from the draft | **Yes** — thought body and redacted metadata (`project` / `branch` / `tags`) are in the reviewer user message | Assembled Trust Gate metadata is scanned before write-back; a supported pattern in reviewer/reasoning/provider/model blocks persist and leaves the draft unchanged |
| `propose_truth` LLM review | Already on disk from the draft | **Yes** — thought body and the selected metadata fields (`project` / `branch` / `tags`; `agent_id` and `metadata.extra` are excluded) are in the reviewer user message | Assembled Trust Gate metadata is scanned before write-back; a supported pattern in reviewer/reasoning/provider/model blocks persist and leaves the draft unchanged |
| Promotion approve | Copy into the permanent namespace; original path removed | Already sent if LLM review ran | |
| Legacy draft that matches the preflight | **Left unchanged** | **Not** sent to the reviewer; **not** copied to a permanent namespace or a supersession successor | Prior persistence is not erased |
| Exports / Rich Views | Approved (and otherwise selected) records already on disk | Local generation | Preflight does not scrub historical files |
Expand Down
2 changes: 1 addition & 1 deletion src/fava_trails/integrations/codev/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -91,4 +91,4 @@ The trust gate reviewer detects artifact type from the scope path:
- `/plans/` in scope → plan checks
- `/reviews/` in scope → review checks

This requires `trail_name` in the redacted metadata (added in fava-trails v0.5.4+).
This requires `trail_name` in the selected metadata fields (added in fava-trails v0.5.4+).
28 changes: 15 additions & 13 deletions src/fava_trails/trust_gate.py
Original file line number Diff line number Diff line change
Expand Up @@ -193,11 +193,12 @@ def describe_trust_gate_egress(
"data_sent": decisions_data_sent,
"data_sent_summary": (
"The full scope-resolved Trust Gate prompt, the full candidate "
"thought content, and the selected redacted metadata (thought_id, "
"source_type, confidence, validation_status; optional "
"trail_name/parent_id/project/branch/tags) are transmitted as "
"structured Decisions state together with the configured Noul "
"question. agent_id and metadata.extra are not sent."
"thought content, and the complete selected metadata fields "
"(thought_id, source_type, confidence, validation_status; "
"optional trail_name/parent_id/project/branch/tags) are "
"transmitted as structured Decisions state together with the "
"configured Noul question. agent_id and metadata.extra are "
"excluded and never sent."
),
"cloud_fallback": False,
"rejection_happens_after_transmission": True,
Expand Down Expand Up @@ -234,8 +235,8 @@ def describe_trust_gate_egress(

data_sent = [
"full candidate thought content (markdown body)",
"redacted metadata: thought_id, source_type, confidence, validation_status",
"optional redacted metadata: trail_name, parent_id, project, branch, tags",
"selected metadata fields sent: thought_id, source_type, confidence, validation_status",
"optional selected metadata fields sent: trail_name, parent_id, project, branch, tags",
]
notice = {
"policy": config.trust_gate,
Expand All @@ -246,10 +247,10 @@ def describe_trust_gate_egress(
"credential_source": trust_gate_credential_description(config),
"data_sent": data_sent,
"data_sent_summary": (
"Candidate thought content plus selected redacted metadata "
"(thought_id, source_type, confidence, validation_status; optional "
"trail_name/parent_id/project/branch/tags). agent_id and metadata.extra "
"are not sent."
"Candidate thought content plus the complete selected metadata "
"fields (thought_id, source_type, confidence, validation_status; "
"optional trail_name/parent_id/project/branch/tags). agent_id and "
"metadata.extra are excluded and never sent."
),
"cloud_fallback": False,
"rejection_happens_after_transmission": True,
Expand Down Expand Up @@ -413,9 +414,10 @@ def prompt_count(self) -> int:


def _redact_metadata(record: ThoughtRecord, *, trail_name: str | None = None) -> dict:
"""Redact sensitive fields from thought metadata before sending to OpenRouter.
"""Select the metadata fields transmitted for review.

Strips: agent_id, metadata.extra, and any fields marked sensitive.
The selected fields are sent in full (their values are not redacted);
agent_id and metadata.extra are excluded entirely.
Includes trail_name (scope path) when provided — not sensitive, enables
scope-based artifact type detection (e.g. /specs/, /plans/, /reviews/).
"""
Expand Down
2 changes: 1 addition & 1 deletion tests/test_trust_gate.py
Original file line number Diff line number Diff line change
Expand Up @@ -353,7 +353,7 @@ def test_redaction_omits_trail_name_when_none(sample_thought):


def test_build_review_payload_includes_trail_name(sample_thought):
"""_build_review_payload passes trail_name through to redacted metadata."""
"""_build_review_payload passes trail_name through to the selected metadata fields."""
trail = "codev-artifacts/Org/Repo/plans/17-hooks"
_, user_msg = _build_review_payload("prompt", sample_thought, trail_name=trail)
assert "trail_name" in user_msg
Expand Down
8 changes: 6 additions & 2 deletions tests/test_trust_gate_decisions.py
Original file line number Diff line number Diff line change
Expand Up @@ -753,8 +753,12 @@ def test_decisions_egress_notice_enumerates_prompt_body_metadata_and_question():
assert "thought_id, source_type, confidence, validation_status" in data_sent
assert "trail_name, parent_id, project, branch, tags" in data_sent
assert "Noul question" in data_sent
# The summary names the exclusions explicitly.
assert "agent_id and metadata.extra are not sent" in notice["data_sent_summary"]
# The summary names the exclusions explicitly and describes the metadata
# as selected fields sent in full — never as "redacted" (issue #125).
summary = notice["data_sent_summary"]
assert "agent_id and metadata.extra are excluded and never sent" in summary
assert "selected metadata fields" in summary
assert "redacted metadata" not in summary.lower()

# Secret-free: credential source is a name only.
assert notice["credential_source"] == "DECISIONS_API_KEY"
Expand Down
38 changes: 37 additions & 1 deletion tests/test_trust_gate_egress.py
Original file line number Diff line number Diff line change
Expand Up @@ -60,7 +60,15 @@ def test_describe_openrouter_default_is_remote_egress_without_secrets():
assert notice["destination_kind"] == "remote_provider"
assert "openrouter" in notice["destination"].lower()
assert "candidate thought content" in notice["data_sent_summary"].lower()
assert "redacted metadata" in notice["data_sent_summary"].lower()
# Regression (issue #125): the disclosure must describe the metadata as
# selected fields sent in full — never as "redacted metadata".
summary = notice["data_sent_summary"].lower()
assert "selected metadata" in summary
assert "redacted metadata" not in summary
assert "redacted" not in " ".join(notice["data_sent"]).lower()
# The exclusion must be stated, not implied by "redacted".
assert "agent_id" in summary
assert "metadata.extra" in summary
assert notice["rejection_happens_after_transmission"] is True
assert notice["cloud_fallback"] is False
text = format_trust_gate_egress_notice(notice)
Expand All @@ -73,6 +81,34 @@ def test_describe_openrouter_default_is_remote_egress_without_secrets():
assert "sk-" not in blob


def test_describe_metadata_disclosure_never_says_redacted_for_either_policy():
"""Regression (issue #125): full prompt, full body, complete selected
metadata fields are sent; agent_id and metadata.extra are excluded.

'Redacted' must not describe the transmitted metadata for any policy.
"""
llm_notice = describe_trust_gate_egress(GlobalConfig())
decisions_notice = describe_trust_gate_egress(
GlobalConfig(
trust_gate="decisions",
trust_gate_decisions_config={
"trust_gate_noul_question": "Acceptance probe?",
"trust_gate_noul_threshold": 0.5,
},
)
)
for notice in (llm_notice, decisions_notice):
blob = json.dumps(notice).lower()
assert "redacted metadata" not in blob
# data_sent lines name the selected fields; none claims redaction.
for line in notice["data_sent"]:
assert "redacted" not in line.lower()
summary = notice["data_sent_summary"].lower()
assert "selected metadata" in summary
assert "agent_id" in summary
assert "metadata.extra" in summary


def test_describe_local_endpoint_marks_local_destination():
notice = describe_trust_gate_egress(
GlobalConfig(
Expand Down
Loading