Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
25 changes: 25 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,31 @@ All notable changes to FAVA Trails are documented here.

## Unreleased

## [0.8.0] — 2026-10-02

**Jev at the boundary between agent output and shared memory.** FAVA Trails
now uses TypeSafe Jev through OpenRouter Decisions by default for new installs.
A typed yes/no question produces a probability; the configured policy decides
whether a draft becomes an approved record. The model, score, threshold, and
review time stay with the record so operators can inspect the decision later.

The shipped `0.45` threshold was selected on a 53-case private replay: 23 of
30 approval-labeled cases accepted, all 23 rejection-labeled cases rejected,
and seven approval-labeled cases rejected. These are reference-label results,
not a guarantee of factual accuracy. See the [calibration report](https://github.com/MachineWisdomAI/fava-trails/issues/126#issuecomment-5894849724)
and [0.8.0 release notes](docs/releases/v0.8.0.md).

### Upgrade

- Existing data repositories and explicit machine settings retain their
selected policy. Set `trust_gate: decisions`, `trust_gate_provider: openrouter`,
and `trust_gate_model: "~typesafe/jev-latest"` to opt an existing installation
into Jev; see [configuration and upgrade guidance](docs/releases/v0.8.0.md#upgrade).
- Jev failures stop promotion; there is no automatic provider fallback.
`llm-oneshot` and explicit operator human approval remain available.
- Each authoring MCP process still needs its own `FAVA_TRAILS_AGENT_ID`.
Identity isolation shipped in 0.7.0 and is not new in this release.

### Fixed
- Trust Gate data-egress disclosure no longer describes the transmitted metadata as "redacted": doctor, the MCP startup notice, and `trust_gate_egress` now state that the full scope-resolved prompt, the full candidate body, and the complete selected metadata fields are sent, with `agent_id` and `metadata.extra` excluded. Docs updated to match; regression coverage added for both `llm-oneshot` and `decisions` notices. Addresses issue #125.

Expand Down
17 changes: 17 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,23 @@ FAVA Trails helps agents share reviewed decisions across sessions and environmen

[Get started](#install) · [Product website](https://fava-trails.org/) · [Two-agent case study](https://fava-trails.org/case-study/) · [Machine Wisdom AI's engineering writing](https://machine-wisdom.ai/writing/)

## New in 0.8: Jev at the memory boundary

An agent can write a plausible conclusion in seconds. Deciding whether the next
agent should inherit it deserves a separate step. FAVA Trails 0.8 uses
[TypeSafe Jev through OpenRouter Decisions](https://openrouter.ai/blog/insights/what-is-jev/)
as its default Trust Gate for new installs: one typed question, an explicit
approval threshold, and a review record you can inspect.

Jev scores the candidate against the supplied memory policy. FAVA applies the
threshold, records the returned model and probability, and makes approved
records available through governed recall. Failed reviews stop promotion.
Your records remain versioned Markdown in your own repository.

[Read the 0.8.0 release notes](docs/releases/v0.8.0.md) for the calibration
results, existing-install configuration, and data-egress details. Review is a
quality control step; it does not establish that a claim is true.

## Why we built FAVA Trails

An observation from one agent can become an assumption for the next. Machine Wisdom AI built FAVA Trails to give that transition an explicit review boundary: saving a thought does not make it accepted knowledge, and an approved correction preserves the record it replaces. The result is a shared history that agents can query and operators can inspect.
Expand Down
90 changes: 90 additions & 0 deletions docs/releases/v0.8.0-announcement.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,90 @@
# Give your agents a memory worth inheriting

*FAVA Trails 0.8.0 brings Jev decision models to shared agent memory.*

An agent's conclusion can outlive the session that produced it. Once it enters
shared memory, another agent may use it to choose an architecture, diagnose a
failure, or continue work days later. That makes the admission decision worth
paying attention to.

**FAVA Trails 0.8.0 puts TypeSafe Jev at that boundary.** The new default Trust
Gate uses OpenRouter Decisions to score whether a candidate deserves to become
durable institutional memory. FAVA applies an explicit threshold and preserves
the result with the record.

The question is straightforward:

> Should this candidate be promoted as durable, high-quality institutional memory under the supplied Trust Gate policy?

Jev returns a probability. FAVA turns that score into an approval or rejection.
The returned model, probability, configured threshold, and review timestamp
become part of the record's provenance. Operators can inspect what happened;
the next agent's default recall sees approved current records.

## A decision model doing a decision job

Jev's appeal here is the interface. The admission step needs a structured
decision, and the Decisions API supplies one directly. FAVA asks a typed Noul
question—the API's yes/no probability type—and applies its policy to the answer.

There is no generated explanation to parse. Jev does not provide one, and FAVA
does not manufacture one. If authentication fails, the service times out, or the
response is invalid, promotion stops. The operator's selected backend remains
the selected backend.

This is a concrete use for decision models in an agent workflow: deciding which
records enter shared memory. The admission policy remains visible and editable
by the operator.

## Evidence behind the default

We selected the `0.45` threshold using a private replay of 53 candidates. Jev
accepted 23 of 30 approval-labeled records and rejected all 23 rejection-labeled
records. Seven approval-labeled records were rejected.

That is zero observed false approvals on this calibration corpus. It is also
a conservative tradeoff: useful records can be missed. The replay informed the
threshold, so these results are calibration evidence rather than an independent
benchmark. Historical review outcomes are reference labels, not ground truth.
The [aggregate report](https://github.com/MachineWisdomAI/fava-trails/issues/126#issuecomment-5894849724)
includes the method and disagreements; the private corpus is available upon
request.

The score measures the model's judgment against the supplied policy. It does
not certify the truth of the candidate's claims.

## Your records remain yours

FAVA Trails is an open-source agent memory system by Machine Wisdom AI. Agents
save drafts, submit them for review, and recall approved records through MCP.
Thoughts remain Markdown files with YAML frontmatter in a versioned,
Git-compatible repository the operator controls. Approved corrections preserve
the lineage of the records they replace.

The Jev integration fits into that existing workflow. It gives the promotion
step a purpose-built decision interface while preserving the readable record
and its history.

Provider choice matters. Hosted review transmits the candidate, policy prompt,
question, and selected metadata before the verdict exists—even when the verdict
is rejection. Explicit `llm-oneshot` review and operator human approval remain
available. Existing installations keep their selected policy until configured
otherwise.

## Try it on a real decision trail

Once 0.8.0 is published, install the release and confirm the loaded runtime:

```bash
pip install --upgrade 'fava-trails==0.8.0'
fava-trails version
```

Restart your MCP connection and follow the [release configuration guide](v0.8.0.md#upgrade).
Use a dedicated agent identity for each authoring MCP process. Start with one
project whose decisions already pass between agents, then inspect what gets
approved, what gets rejected, and whether the records help the next session.

[FAVA Trails](https://fava-trails.org/)
· [Source and installation](https://github.com/MachineWisdomAI/fava-trails)
· [Machine Wisdom AI](https://machine-wisdom.ai/)
107 changes: 107 additions & 0 deletions docs/releases/v0.8.0-launch-kit.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
# FAVA Trails 0.8.0 launch copy

These are announcement drafts for owner publication after the gated release
is verified. The formal positioning workshop remains incomplete; these drafts
describe shipped behavior and calibration evidence, not validated buyer claims.

- Release: https://github.com/MachineWisdomAI/fava-trails/releases/tag/v0.8.0
- Package: https://pypi.org/project/fava-trails/0.8.0/
- Product: https://fava-trails.org/
- Long-form announcement: [Give your agents a memory worth inheriting](v0.8.0-announcement.md)
- Technical detail: [Release notes](v0.8.0.md)

## LinkedIn

Your next agent will inherit today's conclusions. Which ones deserve to stick?

FAVA Trails 0.8.0 brings TypeSafe Jev to that decision.

The new default Trust Gate uses OpenRouter Decisions to score a candidate
against the project's memory policy. FAVA applies an explicit approval threshold
and stores the model, probability, threshold, and review time with the record.

Agents write drafts. Approved records become shared memory. Corrections retain
their history. Everything stays in a versioned Markdown repository you control,
with MCP tools exposing the workflow.

On our 53-case calibration replay, Jev accepted 23 of 30 approval-labeled records
and rejected all 23 rejection-labeled records at threshold 0.45. Seven
approval-labeled records were rejected. Those labels are reference judgments,
not ground truth; the result supports a conservative default rather than a
claim that the system verifies facts.

I like decision models most when there is a concrete decision to make. Here,
it is what the next agent gets to remember.

FAVA Trails is open source, built by Machine Wisdom AI.

Release and installation:
https://github.com/MachineWisdomAI/fava-trails/releases/tag/v0.8.0

## X thread

### 1

FAVA Trails 0.8.0 brings TypeSafe Jev to shared agent memory.

One typed question. An explicit approval threshold. A review record you can inspect.

Give the next agent a reviewed conclusion to inherit.
https://github.com/MachineWisdomAI/fava-trails/releases/tag/v0.8.0

### 2

Jev scores whether a draft deserves to become durable memory. FAVA applies
the threshold and records the returned model, probability, and review time.

Approved records enter default recall. Failed reviews stop promotion.

### 3

At 0.45, our 53-case calibration replay accepted 23/30 approval-labeled records
and rejected all 23 rejection-labeled records. Seven approval cases were missed.

Reference labels, not ground truth. Review quality control, not a fact guarantee.

### 4

Records remain versioned Markdown in your own repo. Agents use MCP to save
drafts, request review, and recall approved records.

Hosted review sends candidate content to the selected provider. Existing
installations retain their policy until configured otherwise.

## Hacker News submission

**Title:** FAVA Trails 0.8: Jev reviews what enters shared agent memory

**URL:** https://github.com/MachineWisdomAI/fava-trails/releases/tag/v0.8.0

**Author comment draft:**

I built FAVA Trails to keep draft observations separate from the records another
agent inherits. It's an MCP server backed by versioned Markdown in a
Git-compatible repository.

0.8 adds TypeSafe Jev through OpenRouter Decisions as the new-install default
reviewer. It asks one typed yes/no probability question, applies an operator-set
threshold, and stores the returned model, score, and threshold with the record.
It stops promotion on review errors and does not invent a model explanation.

The 0.45 default came from a 53-case calibration replay: 23 accepted approval
labels, zero accepted rejection labels, and seven missed approval labels.
The published report explains the method and its limitations. This is a memory
quality gate, not independent fact verification or complete secret detection.

Hosted review transmits candidate content. Explicit chat-model review and
operator human approval remain available. I'd be interested in concrete cases
where a record is useful for the current agent but should not survive into
shared institutional memory.

## Publication checks

Verify the GitHub release is public and PyPI serves 0.8.0 before posting these
drafts. Keep calibration denominators and qualifications with numerical claims.
For the long-form article, replace its conditional publication wording only
after release verification. No social account posting or customer outreach is
performed by creating this kit.
125 changes: 125 additions & 0 deletions docs/releases/v0.8.0.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,125 @@
# FAVA Trails 0.8.0 — Jev at the memory boundary

**Give the next agent a reviewed record to work from.**

FAVA Trails 0.8.0 brings TypeSafe Jev into the point where a draft becomes shared
agent memory. New installs use OpenRouter Decisions by default. The review asks
one typed yes/no question, returns a probability, and applies an explicit
approval threshold. FAVA keeps the model, score, threshold, and review timestamp
with the record.

Agents already produce plenty of text. The useful question is which conclusions
deserve to survive the session and influence the next agent. This release gives
that choice a decision-model interface and an inspectable policy.

## A decision you can inspect

The default question is:

> Should this candidate be promoted as durable, high-quality institutional memory under the supplied Trust Gate policy?

The configured threshold is `0.45`. At or above it, FAVA approves the record;
below it, FAVA rejects promotion. Authentication errors, connection failures,
timeouts, and invalid responses stop promotion. There is no automatic fallback
to another provider. Jev does not return an explanation, and FAVA does not
invent one.

This builds on FAVA's existing lifecycle: agents author drafts, submit them for
review, and recall approved current records. Approved corrections preserve
lineage. Markdown and YAML frontmatter stay in a Git-compatible repository
the operator controls, with MCP tools exposing the workflow.

## How we selected the threshold

The [published calibration report](https://github.com/MachineWisdomAI/fava-trails/issues/126#issuecomment-5894849724)
covers a private 53-case replay with 30 approval reference labels and 23 rejection
reference labels. At `0.45`, Jev accepted 23 approval-labeled cases, rejected all
23 rejection-labeled cases, and rejected seven approval-labeled cases. The
returned model was `typesafe/jev-1.13-20260917`.

That is zero observed false approvals on this replay, with 76.7% recall against
the approval labels. The same replay informed threshold selection, so this is
calibration evidence rather than an independent held-out benchmark. Historical
review labels were reference labels, not verified ground truth. The private
corpus is available upon request.

The practical tradeoff is conservative promotion: some useful records may need
operator review. A high score does not establish factual truth, and a well-formed
decision does not make an agent's underlying claim correct.

## Upgrade

After the gated release is published:

```bash
pip install --upgrade 'fava-trails==0.8.0'
fava-trails version
```

Restart the MCP connection after upgrading. Check the environment the client
actually launches: a wrapper or source-runtime selector can still load an older
checkout after a package upgrade.

New installs default to Jev. Existing data-repository configuration and explicit
machine configuration remain in force. To select Jev on an existing install,
set the following Trust Gate fields in your existing machine configuration
(`~/.config/fava-trails/config.yaml`, or the corresponding `XDG_CONFIG_HOME`
location), preserving other settings:

```yaml
trust_gate: decisions
trust_gate_provider: openrouter
trust_gate_model: "~typesafe/jev-latest"
trust_gate_decisions_config:
trust_gate_noul_question: "Should this candidate be promoted as durable, high-quality institutional memory under the supplied Trust Gate policy?"
trust_gate_noul_threshold: 0.45
```

If that installation has a local `trust_gate_api_base` override, remove it when
switching to hosted OpenRouter. Use the supported OpenRouter credential setup
in the [setup guide](../../AGENTS_SETUP_INSTRUCTIONS.md), then run
`fava-trails doctor` to inspect the effective destination, model, and credential
source without exposing secrets. The moving model alias may resolve to a newer
Jev version; each review records the returned model identity.

Set a stable `FAVA_TRAILS_AGENT_ID` separately in each dedicated authoring
client's MCP launch environment, such as `codex-cli` or `claude-code`. This is a
process identity, not a machine-wide identity. A shared gateway has one identity
boundary. Operator access belongs on a separate operator-controlled endpoint.
See [governed recall](../governed-recall.md).

## Provider choice and privacy

Hosted Decisions sends the full scope-resolved prompt, candidate content,
selected metadata fields, and the question to the configured destination before
a verdict exists. `agent_id` and `metadata.extra` are excluded. Rejected content
has still been transmitted. The existing secret preflight detects specific
credential patterns; it does not provide complete data-loss prevention.

Explicit `llm-oneshot` review and operator human approval remain supported.
The release also includes an explicit local Unsloth Laya Decisions transport.
Its evaluated checkpoint did not separate our approval and rejection cases
well enough for automatic promotion; we retained Jev as the default and deferred
further local-model adoption. Endpoint compatibility alone does not establish
review quality. See [the local evaluation decision](https://github.com/MachineWisdomAI/fava-trails/issues/129#issuecomment-5901330365).

## Included changes

- [Decisions review and durable provenance](https://github.com/MachineWisdomAI/fava-trails/pull/128).
- [Calibrated Jev defaults](https://github.com/MachineWisdomAI/fava-trails/pull/134).
- [Accurate disclosure of transmitted metadata](https://github.com/MachineWisdomAI/fava-trails/pull/130).
- [Explicit local Decisions transport](https://github.com/MachineWisdomAI/fava-trails/pull/135).
- [Machine Wisdom AI attribution](https://github.com/MachineWisdomAI/fava-trails/pull/122)
and [automatic version updates](https://github.com/MachineWisdomAI/fava-trails/pull/136).

## Release verification

Publication uses the existing gated workflow. It validates the exact wheel and
sdist, installed MCP behavior, native registration, identity isolation, and
upgrade from 0.6.0. It publishes those artifacts to PyPI, compares their published
SHA-256 hashes, and then makes the GitHub release public. The release's attached
`candidate-SHA256SUMS` identifies the published artifacts and tagged source.

[Get FAVA Trails](https://github.com/MachineWisdomAI/fava-trails)
· [Product website](https://fava-trails.org/)
· [Built by Machine Wisdom AI](https://machine-wisdom.ai/)
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[project]
name = "fava-trails"
version = "0.7.1"
version = "0.8.0"
description = "FAVA Trails 🫛👣 — Federated Agents Versioned Audit Trail. VCS-backed memory for AI agents via MCP."
readme = "README.md"
requires-python = ">=3.11"
Expand Down
2 changes: 1 addition & 1 deletion src/fava_trails/__init__.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,4 +10,4 @@
try:
__version__ = version("fava-trails")
except PackageNotFoundError: # pragma: no cover - source tree without install metadata
__version__ = "0.7.1"
__version__ = "0.8.0"
Loading
Loading