An agent and memory harness for DittoBench, the benchmark on Bittensor subnet 118 (SN118). Miners run an agent that the validator probes with tool-calling and memory-recall cases. You earn by being more correct than other miners. Latency is reported but not scored. A case that exceeds its per-case timeout scores 0.
The kit is a working baseline plus the full local eval loop (tool calling + memory + speed) running locally against an embedded Turso (SQLite-family) database with native vector search inside the ditto-harness crate.
It mirrors Ditto's production memory retrieval pipeline and ships the real ranking models as weights:
- Vector candidate pool over the seeded memories (cosine on 768-dim embeddings).
- Composite scoring (V2): 7 signals (semantic, linear + exponential recency,
subject frequency, subject semantic match, session continuity, neighbor density)
fused by weights from a weight-predictor MLP (
fixtures/models/mlp-weights.bin, the production architecture retrained on embeddinggemma, which predicts the fusion weights + scale from the query embedding + 17 aux features). - Cross-encoder rerank: a TinyBERT-L2 cross-encoder
(
fixtures/models/cross-encoder.onnx, ONNX viaort) reranks the top-20 pool and fuses with composite rank via Reciprocal Rank Fusion (k=60, ceWeight=0.7).
It also ships a self-contained seed user: a coherent slice of LongMemEval (a public long-term-memory QA benchmark) with subjects already synced.
There are two repositories, and it helps to be exact about the split.
- dittobench-starter-kit (this repo) is the crate you submit. You edit it, then
run
cargo run -- submitto package the whole crate as a tarball. The validator builds that tarball and scores it. This kit is both the local test rig and your submission. - ditto-harness is a dependency, not a copy inside this repo. It is Ditto's
production memory and agent engine, pinned to one commit in
Cargo.toml. Cargo downloads it when you build. No harness source is checked into this kit, so there is nothing here to edit inside it, and you do not submit it.
The dividing line: ditto-harness knows nothing about DittoBench, the validator, or scoring. It is a generic memory and agent library, the same one Ditto runs in production. Every benchmark-aware line is in this kit.
- The engine (ditto-harness): how memories are stored, retrieved, ranked, and reasoned over. The memory store and vector database, the retrieval and ranking pipeline, the agent loop, and the model and embedder clients. It exposes its pieces as slots: it takes an embedder, a ranking-weight predictor, and a reranker, and does not care where they come from.
- The kit (this repo): everything benchmark-specific, plus your work. It fills
those slots with concrete pieces and your weights (
src/baseline.rs), speaks the validator protocol (/health,/seed,/run), ships the tool catalog and the seed user, runs the local practice loop, and is what you customize and submit.
So the loop is: edit this kit, submit this kit. You tune the engine by what you hand it from the kit (your prompt, your reranker, your weights, your retrieval config), including the retrieval algorithm itself, which moves your score most. You rarely edit ditto-harness directly; see Changing the retrieval algorithm for how far the kit reaches and the one case that needs a fork.
src/baseline.rs is the one file most changes start from. It wires the pieces
together (database, embedder, model, retrieval, tools) and is marked with
EXTENSION POINT comments where you plug in your work.
This is what miners optimize, and it is the only thing that moves your score. You
make the harness remember and act better, all from src/baseline.rs:
- Memory retrieval: from a user's history, find the exact past facts that answer a question. This is the harder half of the score.
- Tool use: pick the right tool for a request, and no tool when none fits.
- Orchestration: the prompt and control flow that turn a question into a correct answer.
See How to optimize for the exact knobs, Per-question-type levers for the mechanism behind each scored question type, and Changing the retrieval algorithm for how far the retrieval lever reaches.
Two things are held constant on purpose and never move your score: the model (the validator serves one frozen model and ignores yours) and latency (measured, not scored). See What isn't scored, and why. You do not compete on model choice or hardware.
Why this is the competition: Ditto's product is memory, an assistant that recalls what you told it across sessions. DittoBench scores that exact capability on Ditto's real production retrieval stack, not a stand-in, so a harness that scores higher is a better memory system, and strong work can flow back into the product. The composite score is half tool accuracy, half memory recall.
The loop is: edit this kit, test locally, submit this kit. In order:
- Edit
src/baseline.rs: change the prompt, the retrieval config, the reranker, the ranking weights, or the tools (the levers above). - Test retrieval, fast and offline:
cargo run -- mem-eval --k 10reports recall@k with no LLM (needs only Ollama for embeddings). Run this after any change to retrieval, weights, or the reranker. - Test the full agent:
cargo run -- evaluate(fixed inputs, best for iterating) orcargo run -- practice(rotating inputs). Watch the composite, the per-category tool means, and the slowest cases. Run this after any change to the prompt or tools. - Rehearse against the real validator (optional, recommended before you submit): serve your harness and drive it from the playground Submit tab. This is the only local run with a fresh random dataset per submission, and the only one that exercises Tier B/C seeding. See Hosted BYOK (bring your own key) practice.
- Package:
cargo run -- submitbuildsdittobench-submission.tgzfrom your whole crate. - Go on-chain: register a hotkey on netuid 118 and upload with the eval fee. The validator builds your crate in Docker and scores it under the model lock. See Mining on SN118.
You never leave this repo for any of it. ditto-harness is fetched automatically when you build.
| File | What it is |
|---|---|
src/baseline.rs |
The agent you optimize. Wires DB + embedder + model + MLP predictor + reranker + harness. |
src/reranker.rs |
ONNX cross-encoder reranker, the production rerank stage, 1:1. |
src/seed.rs |
Loads the bundled LongMemEval seed user into the vector DB. |
src/protocol.rs |
The validator HTTP wire contract (see PROTOCOL.md). |
src/catalog.rs |
The Ditto tool catalog presented per case. |
src/datagen.rs |
Deterministic-per-seed dataset generator. |
src/scorer.rs |
Local score report (tool accuracy + memory + latency). |
src/bin/dittobench-miner.rs |
CLI: serve, playground, seed-user, mem-eval, evaluate, practice, submit. |
fixtures/seed-user/ |
The seed user: pairs + pre-synced subjects + subject graph + LongMemEval questions. |
fixtures/models/ |
Shipped weights: mlp-weights.bin (217K-param MLP) + cross-encoder.onnx (TinyBERT-L2 INT8) + BERT vocab. |
scripts/build-seed-user.py |
Regenerates the seed-user slice from the LongMemEval fixture (maintainers only, inputs not distributed). |
SETUP.md is the step-by-step guide for setting up this kit with
ditto-harness(the crate dependency), including Ollama and.env.
# 1. Pick a chat model for LOCAL practice. Default provider is OpenRouter,
# defaulting to qwen/qwen3-32b, the on-chain scored model (matches
# .env.example). On-chain scoring locks inference to Qwen3-32B served in a
# TEE (Trusted Execution Environment; Chutes Qwen/Qwen3-32B-TEE) and overrides
# your model, so the default already matches scoring conditions.
export OPENROUTER_API_KEY=sk-or-...
# (optional) export DITTOBENCH_MODEL=<any OpenRouter model id>
# ...or use Chutes hosted OpenAI-compatible inference (set
# DITTOBENCH_MODEL=Qwen/Qwen3-32B-TEE to practice on the exact scored backend):
# export DITTOBENCH_PROVIDER=chutes
# export CHUTES_API_KEY=cpk_...
# export DITTOBENCH_MODEL=Qwen/Qwen3-32B-TEE # Chutes defaults to deepseek-ai/DeepSeek-V3.2-TEE; override to the scored model
# ...or run fully local with Ollama:
# export DITTOBENCH_PROVIDER=ollama
# export DITTOBENCH_MODEL=qwen2.5:7b
# 2. Embeddings use Ollama's embeddinggemma (768-dim) by default. For memory
# cases you need it running locally:
# ollama serve
# ollama pull embeddinggemma
# 3. Load the seed user (one-time, embeds pairs + subjects), then practice.
cargo run -- seed-user # load the LongMemEval seed user
cargo run -- mem-eval --k 10 # retrieval recall over the seed user (no LLM)
cargo run -- evaluate # FIXED local submission test (static user + same questions)
cargo run -- practice --n 20 # ROTATING random dataset (anti-overfit, like the hosted validator)
# 4. Serve the harness for the validator.
cargo run -- serve --port 8080The interactive playground is a chat UI wired to a production-Ditto agent:
the v2 system prompt + persona + tool-use policy, the model set by
DITTOBENCH_MODEL (.env.example ships qwen/qwen3-32b, the scored model; set
google/gemini-3.1-flash-lite to mirror prod Ditto's model), the full tool
catalog, and real memory retrieval + cross-encoder rerank over the seed user. Action tools
(search_web, create_image, agent jobs, settings, …) return fake-but-plausible
results so you can exercise tool-calling without real integrations. Memory
tools are real and query the seed user.
cp .env.example .env # paste your OPENROUTER_API_KEY into .env
cargo run -- seed-user # one-time: load the dummy seed user
cargo run -- playground # open http://127.0.0.1:8088The UI shows the full tool catalog (every tool's description + JSON schema),
and after each turn a trace of the tool calls (args + fake results) and
the memories retrieved for that query. Try "search the web for…"
(search_web fires) or "how many postcards have I collected?" (memory
retrieval answers with ditto://memory/… citations). The Submit tab scores
your harness against the official hosted validator. See Hosted BYOK practice
below.
evaluate(local, fixed): scores your submission against the same inputs every run: the static seed user, the same bundled LongMemEval questions, and a fixed-seed tool set. Inputs are reproducible and model output is still stochastic.practice(local, rotating): re-rolls prompts per run, but from a small fixed template pool (10 memory facts). It varies wording, not substance, and never exercises the seeding tiers/waves.- Hosted validator: generates a fresh random dataset per submission, like the on-chain SN118 validator, and is the only pre-chain rehearsal of Tier B/C seeding and the real question mix. Drive it from the playground's Submit tab (below). Scoring runs under the locked model, so keep
DITTOBENCH_MODELat the kit default (qwen/qwen3-32b), which on-chain scoring serves in a TEE. A crate submission is built and run under the hard lock; a harness_url submission uses your harness's configured model, so leave the default in place.
Use evaluate to develop.
The hosted validator is available. The playground's Submit tab drives it:
- Serve your harness and expose it publicly so the validator can reach it:
cargo run -- serve --port 8080, then e.g.ngrok http 8080. - Set
DITTOBENCH_HARNESS_URLin.envto the public URL. .env.example ships the officialDITTOBENCH_API_URL. cargo run -- playground→ open the Submit tab and pick a run size. Your harness calls the locked model (Qwen3-32B, the kit default) and yourOPENROUTER_API_KEYpays for that inference; BYOK means your key, not your choice of model. The validator stores no keys. LeaveDITTOBENCH_MODELat the default so practice matches scoring.
Two targets: local (your serve exposed publicly, as above) or crate
(the validator builds your repo from a git URL, which must be publicly
fetchable, and a private fork needs the gh_token path).
seed-user and mem-eval need only Ollama (embeddinggemma). No chat model
or API key is required. mem-eval runs the full
production pipeline (MLP weights + composite V2 + cross-encoder rerank) and
reports recall@k per LongMemEval question type, isolating retrieval quality
from the LLM. Keep the same DITTOBENCH_DB across seed-user and mem-eval.
cargo build and cargo test need no model or embedder, but the first
build needs network (the git dependency fetch, and the ort crate downloads
ONNX Runtime). Rebuilds are offline. Only practice/serve call out
to the model + Ollama at runtime.
The validator calls POST /run with a RunRequest (system prompt, user
input, available tools) and expects a RunResponse (final text, observed tool
calls, token usage, latency). Before memory questions it installs a haystack via
POST /seed. Full shapes in PROTOCOL.md.
Every submission gets a fresh procedural persona universe, and the composite
is 0.5 × tool + 0.5 × memory. The full grading rubric lives in PROTOCOL.md.
BASELINES.md reports what the stock kit scores under the locked
model (the target to beat) and its weakest categories.
Memory is seeded in three tiers (see PROTOCOL.md POST /seed):
A prepared subjects, B raw pairs only (build your own subject index),
C staged waves interleaved with runs (upsert each
wave). The bundled harness reuses the production save_memory path in
seed.rs. Extending it to construct subjects when none are provided (Tier B) is
the highest-value change you can make.
The three levers from What you optimize, in detail. Everything here lives in
src/baseline.rs, marked EXTENSION POINT. The chat model is locked (see What
isn't scored, and why), so the levers are retrieval, the prompt, and tools:
- Retrieval / memory: the production stack is wired and active, including the
weight-predictor MLP, composite V2, and the cross-encoder reranker
(
open_store). Tune it by retraining/swappingfixtures/models/mlp-weights.bin, swapping the cross-encoder ONNX, adjusting the RRFk/ceWeightinreranker.rs, or changingcandidate_pool_size/variant/limits. Measure withmem-eval(recall@k). Memory is the harder half of the composite and retrieval recall is the main bottleneck, so this is the highest-value scored lever. - System prompt: augment the per-case prompt with a tool-use policy and abstention rules so the agent picks the right tool (and no tool when it shouldn't).
- Tools: the baseline registers the per-case tool catalog as stub tools so
the agent can select the right one (what the validator scores). Add real
host
Toolimplementations (WireTool→ your own) to execute tools.
Two things you might expect to tune do not affect your score: the model and latency. Both are held out on purpose.
Model. Every miner is scored on the same frozen model. A scored crate run builds
your image and serves it under the lock: the validator's relay overrides
DITTOBENCH_MODEL and serves Qwen3-32B in a TEE regardless of what baseline.rs
sets (see Mining on SN118), so swapping the model changes only local practice
speed and cost. The benchmark measures the harness (memory, retrieval, agent
orchestration, tool selection), not the model. If model choice were scored, the
board would rank who can afford the strongest frontier model, not who built the
best agent, turning an open-source harness competition into a spend race. One
frozen open-weight model holds that variable constant, so score gaps reflect
harness quality and every miner is scored under identical, attestable conditions:
the TEE proves the locked model actually ran, so no one can quietly swap in a
stronger one. Keep the kit default (qwen/qwen3-32b, the same weights) so
practice tracks scoring.
Latency. Reported as median_ms on the leaderboard, never scored. It measures
hardware and model-provider speed, not harness quality, and it varies with
sandbox load, so two validators would compute different numbers, and the weight
fold requires every validator to derive the same score from the same public
ledger. Speed is bounded instead of ranked: a case that misses its per-case
timeout scores 0, and the tool-efficiency factor penalizes over-calling, so the
efficiency that reflects harness quality is captured without tying the score to
raw hardware speed.
Put your effort into the three levers above.
Run mem-eval after retrieval changes (recall@k, no LLM) and practice after
agent/tool changes (watch composite, per-category tool means, slowest cases).
Every memory question type maps to a concrete mechanism in this kit (or one you can add). Nothing is scored that lacks a lever; the lever is the capability being measured:
| Question type | Miner lever |
|---|---|
| single/multi-session, preference, knowledge-update | retrieval quality: composite signals, the subject index, recency handling (latest value wins) |
| temporal, trajectory, duration | timestamp arithmetic over the seeded pairs' timestamps. Order and elapsed time come from the transcript, not the model's guess |
| assistant-recall | store and index the ASSISTANT turns, not just user turns (the answer only ever appeared in an assistant reply) |
| aggregation, computed | mention counting across sessions; deliberately punishes naive dedup collapse of repeated topics |
| canary | a lexical/exact-match index. Embeddings represent random tokens (VK-… codes) poorly, so semantic-only retrieval misses them: a concrete, winnable gap the stock kit does not attempt |
| injection-resistance | a system-prompt guard: the frozen harness model complies with embedded overrides unless YOUR harness defends; a scored, discriminative surface the stock kit does not attempt |
| isolation | honor user_id scoping (already wired in the kit's store) |
| abstention, DRM lure (a related decoy that tempts a false recall) | confidence gating + the abstain wire flag. Decline when retrieval finds nothing (or finds only someone ELSE's value) instead of fabricating |
| contradiction | read the LATEST stance from memory: some opinions were reversed ("no longer do it") and some were not ("still love it"). Both answers occur under the same question surface, so the signal must come from retrieval |
The two rows that most separate a naive submission from a competitive one are canary (needs a lexical index) and injection (needs a prompt guard): both are scored on every run and the stock kit leaves them on the table.
The kit defaults to local Ollama embeddinggemma (768-dim) for a free,
self-contained loop. To make the ranker work in that space, the shipped MLP is
retrained on embeddinggemma (via the production training pipeline, on
LongMemEval).
On the bundled seed user this lifts retrieval from hit@10 0.90 → 0.96 vs the
Vertex-trained weights. The cross-encoder rerank is embedder-independent (it
scores raw text), so it is identical to production.
If you switch build_embedder to a different embedder, retrain the MLP for
that space. The production training pipeline is not distributed, but the
artifact format is documented (fixtures/models/README.md
and the harness's mlp.rs), so you can train your own with any pipeline and
drop it in. To run the exact production stack, use Vertex text-embedding-005
- the production
model.bin.
Retrieval is the main lever on your score, so this is worth being clear about:
you can change the ranking algorithm itself, and for the most part you do it from
this kit, not by editing ditto-harness. The harness exposes the pipeline as
seams that baseline.rs already wires up:
- Reranker: it is a trait.
baseline.rsbuilds anArc<dyn Reranker>and injects it. Implement your own reranker in yoursrc/and pass it in, or swap the model infixtures/models/cross-encoder.onnx. - Fusion weights: the weight predictor loads your bytes
(
MlpPredictor::load_from_reader), so retrain and drop in your ownfixtures/models/mlp-weights.bin. - Raw retrieval:
store.db()exposes the database directly (candidate search, subjects, raw rows). You can bypass the built-in ranker entirely and score candidates with your own algorithm in your crate. - Candidate pool and variant:
CompositeSearchRequestexposescandidate_pool_size,variant, and limits. - New retrieval capabilities (a lexical index for canary codes, a subject index,
indexing assistant turns): add them in your
src/on top of the store API.
You only fork ditto-harness to edit its built-in composite scorer in place (its
internal signal math and candidate queries), which you can otherwise override or
bypass with the seams above. To fork: point the ditto-harness git URL and
rev in Cargo.toml at your fork. Editing a local clone has no effect on your
submission unless you repoint the dependency, because the build uses the pinned
commit. Your fork must be publicly fetchable for the validator to build it; a
private fork needs the gh_token path shown in the Dockerfile.
cargo run -- submit # packages dittobench-submission.tgz + prints next stepssubmit runs tar -czf dittobench-submission.tgz . (excluding target/,
.git, *.tgz, *.db, *.db-*, .env, .env.*, and it prints the exclusion
list). Never commit or package your .env. It holds your
OPENROUTER_API_KEY, and the tarball is uploaded to the platform. You submit
the entire buildable project, with the Dockerfile at the tarball root:
Dockerfile,Cargo.toml,Cargo.locksrc/, including your editedbaseline.rsand thedittobench-minerserverfixtures/, the ONNX models + seed data your harness loads at runtime
The tarball is capped at 20 MiB. The shipped fixtures/models/ fit well under
that; if you bundle larger weights, keep the packaged tarball within the cap.
You are not submitting src/baseline.rs on its own, and you are not
submitting ditto-harness. ditto-harness is a pinned public git dependency
of your crate. The Docker build fetches it.
The validator builds your tarball in Docker, runs the resulting container, then scores it. A submission is only valid if it keeps this contract intact:
| Must hold | Why |
|---|---|
A Dockerfile at the tarball root |
It's the validator's Docker build context. |
docker build succeeds |
A pre-screen gate rejects submissions that don't build. |
The image serves GET /health, POST /seed, POST /run on :8080 |
The validator drives your harness over these (see PROTOCOL.md). |
POST /run returns a well-formed RunResponse |
The scorer grades tool_calls + final_text. A malformed body scores 0. |
Everything else is yours: baseline.rs holds the scored levers from How to
optimize (system prompt, retrieval knobs, tools) plus the model, which only
affects local practice; then any other src/ file, added crate dependencies,
your own fixtures/models/ weights, even the Dockerfile build steps.
Restructure the crate however you like, as long as docker build . still
produces a container serving that protocol on :8080.
Status: the hosted practice validator and the SN118 leaderboard are live today. The on-chain submission path (
ditto upload, eval fee, scoring, weights) is not yet live, so no competitive scores populate the leaderboard yet. It runs onbench_version2, and this section documents that contract.
-
Registration. You need a hotkey registered on subnet netuid 118 (
btcli subnet register --netuid 118) and TAO for the registration cost plus per-submission eval fees. -
Submission + fee. This kit's
submitonly packages the tarball. The on-chain upload happens throughditto upload(the miner CLI from the ditto-subnet repo), with your registered hotkey. Each upload pays a per-submission eval fee of roughly $5 USD, quoted in TAO at upload time (an oracle sets the exact amount, and the CLI shows the quote before you confirm). The fee is the effective rate limit. -
The runtime contract. Scoring is one-shot from your uploaded tarball. You do not keep a server running. The scorer builds your
Dockerfilein a sandbox, starts yourserveprocess, and injects env at runtime:OPENROUTER_API_KEY(the validator's key pays your harness's inference on on-chain runs, and your own key on hosted BYOK practice),DITTOBENCH_PROVIDER/DITTOBENCH_MODEL, and a freshDITTOBENCH_DBpath. The Docker host gateway is mapped so the defaultOLLAMA_BASE_URLresolves to the scoring host's Ollama, which serves the referenceembeddinggemmaembedder. If you use a different local embedder, bundle it in your image. The container has network egress to model providers (hardened deployments may restrict egress to an allowlist). Your harness must read all model config from env.Baseline::from_envalready does this, so keep that property if you rewrite it. -
Timeouts. 10 s for
/healthto come up, 60 s per/runcall (a case that misses it scores 0), 5 minutes per/seedwave. The table is in PROTOCOL.md. -
Run shape. An on-chain run is
run_size=full: on the order of 50 memory cases + 60 tool cases, with 2 staged seeding waves, a substantial Tier-B raw-pairs share, and a handful of isolation cases across separate user graphs. Exact counts can change withbench_version. -
Economics (king-of-the-hill, winner-take-most). The champion receives ~90% of the miner emission, the next 4 ranked miners split the remaining ~10%, and everyone else earns nothing. A challenger dethrones the champion only by beating its composite by more than a 5% relative margin (plus a statistical uncertainty band when score error bars are available). Weights are recomputed from the public score ledger on every validator sweep. Being 2nd by 4% earns a tail share.
-
bench_version. Only scores from the latest
bench_versioncompete. When it bumps, validators automatically re-score the champion and top tail on the new version, so you don't need to resubmit or re-pay, though your standing can change.bench_version2 is the launch version. -
Lifecycle. After upload your agent goes
uploaded → evaluating → scored, orscreening_failedif the Docker build or/healthfails (fix and resubmit). Scores land on the public score ledger and the SN118 leaderboard, with median latency,bench_version, and per-category stats. -
Anti-gaming. The dataset is procedurally regenerated per run from a fresh seed. Grading is deterministic (no judge to prompt-inject): emitting an embedded injection payload, surfacing another user's value, or naming a distractor value zeroes the case, and those events land in the run's public details. Malformed responses, timeouts, and build failures score 0. Observed tool execution (the validator runs your tool calls against its own mock endpoint) caps unverified self-reported tool calls at 0.5. Beyond per-case grading, the composite carries bounded integrity multipliers: a per-run canary nonce (bounded penalty for an honest miss, hard cap for leaking the decoy), a metamorphic-consistency factor over invariance families, and the tool-efficiency factor. All three are detailed in PROTOCOL.md and are pure functions of the published run details.
-
Originality (duplicate detection). Before scoring, the platform compares your uploaded crate against every other miner's eligible submission across several dimensions: exact bytes, normalized source (comments, whitespace, and formatting stripped, so a reformat/recomment/file-rename does not hide a copy), lexical and AST-structural fingerprints, the prompt/strategy text, and a semantic code-embedding vector. A runtime behavioral signal (your observed tool-call trajectory on a shared dataset seed) is not folded into this comparison today. An exact or trivially repackaged copy of another miner's agent is held for manual review and excluded from the ledger (zero weight) while held. The earlier upload wins by first-seen. In the softer "similar but not identical" band a hold requires agreement across multiple independent signals, so legitimate convergence is not penalized: building on this starter kit and the shared
ditto-harnessdependency is expected, and two miners independently arriving at similar prompts or structure is fine. Detection targets copying another miner's submission; shared use of the public baseline is expected. First-seen protects the original author over a later uploader. Forking a leaked crate and renaming its symbols does not earn. Differentiate on substance: change the prompt, retrieval, or tools (see How to optimize). The model is not a differentiator, since scoring locks every miner to the same frozen model. -
Hardware. The reference stack runs on CPU: Ollama
embeddinggemma(~1 GB), the TinyBERT ONNX reranker, embedded SQLite. 8 GB RAM is sufficient. No GPU is required unless you bundle a local LLM. -
Support. Open a GitHub issue on this repo.
- Do not overfit the local scorer. The local dataset generator is a simplified pool while the validator's persona universe rotates every run. See the blockquote in PROTOCOL.md.
- Arguments weigh as much as selection on-chain. The on-chain
deterministic tool grade is
0.4 name F1 + 0.4 argument F1 + 0.2 trajectory(see PROTOCOL.md). Only the local scorer is name-centric. - Latency is not scored (only reported as
median_ms); a case that misses its per-case timeout scores 0, and the tool-efficiency multiplier bounds over-calling on observed runs. The rationale is in What isn't scored, and why. Measure withpractice. - Memory needs the seed user loaded + Ollama embeddings. Run
seed-userfirst. Ifmem-evalreportsrecall@k: 0.000, see SETUP.md → Troubleshooting.
MIT (LICENSE). The ditto-harness dependency is also MIT-licensed.