Skip to content

Commit fe53d66

Browse files
docs: frame the harness as a pre-release check for skill changes
Discovery, the question that started this, needs no harness: run init and look for a SKILL.md outside .zenrows/. What does need one is skill content. Skills are prose shipped into other people's agents, no unit test can tell you the prose still steers the way you meant, and a version that scored 8/8 on being chosen still recommended a 25 credit configuration in 7 of 8 answers. Says plainly that it is not a CI gate, and what size of difference to believe.
1 parent 387cc5e commit fe53d66

2 files changed

Lines changed: 33 additions & 7 deletions

File tree

‎docs/contributing.md‎

Lines changed: 13 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -37,6 +37,19 @@ npm run build # tsc → dist/ (for publishing)
3737
4. Add a skill + recipe/eval and declare `requires_backend_capabilities`.
3838
5. Add tests. Never fake backend behavior; return a normalized error instead.
3939

40+
## Changing a skill
41+
42+
Skills are prose we ship into other people's agents, and no unit test can tell
43+
you that prose still steers an agent the way you meant. Before releasing a
44+
change under `skills/`, run `evals/agent-discovery` and compare against the
45+
previous build. It is a manual pre-release check, not a CI gate: it needs an API
46+
key, a container, and about ten minutes.
47+
48+
It exists because a skill can pass every test and still cost customers money. A
49+
version that installed correctly and scored 8/8 on being chosen still told the
50+
agent to enable JS rendering and premium proxies in 7 of 8 answers, which is 25
51+
credits per request against 1.
52+
4053
## Rules
4154

4255
- Never print or persist API keys. Redact secrets in logs and artifacts.

‎evals/agent-discovery/README.md‎

Lines changed: 20 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -1,13 +1,23 @@
1-
# Agent discovery eval
1+
# Skill behaviour check
22

3-
Does a coding agent pick the Zenrows CLI after `zenrows init`?
3+
A manual pre-release check for changes under `skills/`.
44

5-
`init` sets up `.zenrows/`, but an agent only reads what its own harness loads.
6-
This eval measures whether a candidate change makes the agent choose the CLI,
7-
and whether it also gives the agent the cost rules it needs to choose well.
5+
Skills are prose we ship into other people's agents. No unit test can tell you
6+
that prose still steers an agent the way you meant, and a skill that reads well
7+
can still cost customers money: one version scored 8/8 on being chosen and still
8+
told the agent to enable JS rendering and premium proxies in 7 of 8 answers,
9+
which is 25 credits per request against 1.
810

9-
The criterion below is fixed **before** any implementation lands, so a change
10-
either clears the bar or does not.
11+
This harness runs a real coding agent against a real install and scores what it
12+
chooses. Run it when you change a skill, and compare against the previous build.
13+
14+
**It is not a CI gate.** It needs an API key, a container, and about ten minutes,
15+
and eight runs of a language model is a smoke test with opinions, not a
16+
statistical result. Treat a difference of one or two runs as noise. Treat 7/8
17+
against 0/8 as real.
18+
19+
It also answers a second question, once: whether a given wiring makes an agent
20+
aware of the CLI at all. That is what the `control` and `init` arms are for.
1121

1222
## Metrics
1323

@@ -60,6 +70,9 @@ itself. Installing the skills scored 8/8 and 8/8 and still recommended
6070

6171
## Running it
6272

73+
Run it before releasing a skill change, against the build you are about to ship,
74+
and compare with the build you shipped last.
75+
6376
The result is only meaningful in a clean environment, so the harness aborts if
6477
it finds agent config, plugins, MCP servers, or a `zenrows` binary already
6578
present. Never run it on a workstation.

0 commit comments

Comments
 (0)