You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 8152ca5
Browse filesBrowse the repository at this point in the historyBrowse files
Copy file name to clipboardExpand all lines: evals/capability/README.md
+17-12Lines changed: 17 additions & 12 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -91,35 +91,40 @@ shell parser.
91
91
92
92
## Prerequisites
93
93
94
-
-Configured provider (same as interactive `corbits`)
94
+
-A configured provider matching the CLI `--provider` / `--model` (or `--matrix` cells)
95
95
- Network access for inference
96
96
- Bun
97
97
98
98
Evals default to `--dangerously-skip-permissions` so the agent can write without a human at the gate. Override with `--ask-permissions` if you want the non-interactive deny path.
99
99
100
+
## Provider/model
101
+
102
+
CLI `--provider <name>` and `--model <id>` are required for every run, including `--dry-run`. Alternatively, pass `--matrix` with complete `provider:model` cells. Local `.corbits/settings.json` is never the implicit eval target.
103
+
100
104
## Run
101
105
102
106
```bash
103
-
# All cases with the configured default provider/model
104
-
bun run eval:capability
107
+
# All cases — explicit provider/model required
108
+
bun run eval:capability -- --provider <name> --model <id>
105
109
106
110
# One case, explicit model
107
-
bun run eval:capability -- --case simple-health --provider xai/thegreataxios --model grok-4.5
111
+
bun run eval:capability -- --case simple-health --provider xai --model grok-4.5
@@ -129,8 +134,8 @@ bun run eval:capability -- --repeats 5 \
129
134
Any change intended to shift agent behavior (prompts, directors, tools) is
130
135
confirmed here, not by anecdote:
131
136
132
-
1. Run the suite with `--repeats 5` (repeats smooth model variance; a single
133
-
run of a bait case proves nothing).
137
+
1. Run the suite with `--provider` / `--model` (or `--matrix`) and `--repeats 5`
138
+
(repeats smooth model variance; a single run of a bait case proves nothing).
134
139
2. Compare against the frozen baseline
135
140
(`evals/capability/results/baseline-0286.json`) with `--baseline`.
136
141
3. Read the verdicts: any pass-rate change per cell is significant; behavior
@@ -148,8 +153,8 @@ Flags:
148
153
| Flag | Meaning |
149
154
|------|---------|
150
155
|`--case <id\|all>`| Case id or `all` (default) |
151
-
|`--provider` / `--model`| Single-variant override via `loadConfig`|
152
-
|`--matrix <cells>`| Multi-variant: `p:m,p2:m2` or `label=p:m` (comma-separated) |
156
+
|`--provider <name>` / `--model <id>`|Required unless `--matrix`. Single-variant via `loadConfig`. Not inferred from local settings|
157
+
|`--matrix <cells>`|Alternative to `--provider`/`--model`. Multi-variant: `p:m,p2:m2` or `label=p:m` (comma-separated). Every cell must include both sides|
|`--repeats <n>`| Runs per case×variant cell (default `1`; gate runs use `5`, baseline freezes `3`). Results record every repeat plus per-cell aggregates |
161
-
|`--dry-run`| Load cases × variants and print plan; no inference |
166
+
|`--dry-run`| Load cases × variants and print plan; no inference. Still requires `--provider`/`--model` or `--matrix`|
0 commit comments