Skip to content

Commit 8152ca5

Browse files
committed
Require provider and model flags for all eval runners
Local settings must not pick the model. Capability eval and SWE dry-run now require --provider/--model the same as a real SWE run.
1 parent beb9315 commit 8152ca5

9 files changed

Lines changed: 157 additions & 42 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 4 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -78,6 +78,10 @@ parallel copies under `docs/` or `scripts/notes/`. At cut time: rename
7878

7979
### Evals
8080

81+
- **Eval runners require an explicit model pair.** `eval:capability` and
82+
`eval:public-swe-one` take `--provider` / `--model` (capability also
83+
accepts `--matrix` with complete cells) so local `.corbits/settings.json`
84+
is not the implicit target.
8185
- **Capability eval records `task` tool calls.** `taskToolCallCount` is derived
8286
from the turn stream (informational). Older result files without the field
8387
default from `toolCallsByName.task` so the frozen baseline still parses.

‎evals/capability/README.md‎

Lines changed: 17 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -91,35 +91,40 @@ shell parser.
9191

9292
## Prerequisites
9393

94-
- Configured provider (same as interactive `corbits`)
94+
- A configured provider matching the CLI `--provider` / `--model` (or `--matrix` cells)
9595
- Network access for inference
9696
- Bun
9797

9898
Evals default to `--dangerously-skip-permissions` so the agent can write without a human at the gate. Override with `--ask-permissions` if you want the non-interactive deny path.
9999

100+
## Provider/model
101+
102+
CLI `--provider <name>` and `--model <id>` are required for every run, including `--dry-run`. Alternatively, pass `--matrix` with complete `provider:model` cells. Local `.corbits/settings.json` is never the implicit eval target.
103+
100104
## Run
101105

102106
```bash
103-
# All cases with the configured default provider/model
104-
bun run eval:capability
107+
# All cases — explicit provider/model required
108+
bun run eval:capability -- --provider <name> --model <id>
105109

106110
# One case, explicit model
107-
bun run eval:capability -- --case simple-health --provider xai/thegreataxios --model grok-4.5
111+
bun run eval:capability -- --case simple-health --provider xai --model grok-4.5
108112

109113
# Multi-model matrix (cases × variants)
110114
bun run eval:capability -- \
111-
--matrix "xai/thegreataxios:grok-4.5,openai:gpt-4.1" \
115+
--matrix "xai:grok-4.5,openai:gpt-4.1" \
112116
--out evals/capability/results/matrix.json
113117

114118
# Labeled variants
115119
bun run eval:capability -- --matrix "fast=xai:grok-4.5,strong=openai:gpt-4.1"
116120

117121
# Baseline improve/regress (keys by variantId::caseId)
118-
bun run eval:capability -- --out evals/capability/results/run2.json \
122+
bun run eval:capability -- --provider <name> --model <id> \
123+
--out evals/capability/results/run2.json \
119124
--baseline evals/capability/results/run1.json
120125

121126
# Gate run: 5 repeats per cell against the frozen baseline
122-
bun run eval:capability -- --repeats 5 \
127+
bun run eval:capability -- --provider <name> --model <id> --repeats 5 \
123128
--out evals/capability/results/candidate.json \
124129
--baseline evals/capability/results/baseline-0286.json
125130
```
@@ -129,8 +134,8 @@ bun run eval:capability -- --repeats 5 \
129134
Any change intended to shift agent behavior (prompts, directors, tools) is
130135
confirmed here, not by anecdote:
131136

132-
1. Run the suite with `--repeats 5` (repeats smooth model variance; a single
133-
run of a bait case proves nothing).
137+
1. Run the suite with `--provider` / `--model` (or `--matrix`) and `--repeats 5`
138+
(repeats smooth model variance; a single run of a bait case proves nothing).
134139
2. Compare against the frozen baseline
135140
(`evals/capability/results/baseline-0286.json`) with `--baseline`.
136141
3. Read the verdicts: any pass-rate change per cell is significant; behavior
@@ -148,8 +153,8 @@ Flags:
148153
| Flag | Meaning |
149154
|------|---------|
150155
| `--case <id\|all>` | Case id or `all` (default) |
151-
| `--provider` / `--model` | Single-variant override via `loadConfig` |
152-
| `--matrix <cells>` | Multi-variant: `p:m,p2:m2` or `label=p:m` (comma-separated) |
156+
| `--provider <name>` / `--model <id>` | Required unless `--matrix`. Single-variant via `loadConfig`. Not inferred from local settings |
157+
| `--matrix <cells>` | Alternative to `--provider`/`--model`. Multi-variant: `p:m,p2:m2` or `label=p:m` (comma-separated). Every cell must include both sides |
153158
| `--config <path>` | Settings file override (CI injection) |
154159
| `--out <path>` | Write machine-readable results JSON |
155160
| `--baseline <path>` | Compare this run to a prior results file (improve/regress + metric deltas) |
@@ -158,7 +163,7 @@ Flags:
158163
| `--agent-timeout-ms <n>` | Wall-clock limit for `runExec` (default `600000`, env `CORBITS_EVAL_AGENT_TIMEOUT_MS`) |
159164
| `--verify-timeout-ms <n>` | Wall-clock limit for `verify.sh` (default `120000`, env `CORBITS_EVAL_VERIFY_TIMEOUT_MS`) |
160165
| `--repeats <n>` | Runs per case×variant cell (default `1`; gate runs use `5`, baseline freezes `3`). Results record every repeat plus per-cell aggregates |
161-
| `--dry-run` | Load cases × variants and print plan; no inference |
166+
| `--dry-run` | Load cases × variants and print plan; no inference. Still requires `--provider`/`--model` or `--matrix` |
162167

163168
## Case format
164169

‎evals/capability/lib.test.ts‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -309,6 +309,16 @@ describe("parseMatrix", () => {
309309
expect(v[0]!.provider).toBe("xai");
310310
expect(v[0]!.model).toBe("thegreataxios/grok-4.5");
311311
});
312+
313+
test("rejects incomplete cells", () => {
314+
expect(() => parseMatrix("xai:", {})).toThrow(/both provider and model/);
315+
expect(() => parseMatrix(":grok-4.5", {})).toThrow(/both provider and model/);
316+
});
317+
318+
test("fills omitted cell side from --provider/--model defaults", () => {
319+
const v = parseMatrix("xai:", { model: "grok-4.5" });
320+
expect(v[0]).toEqual({ id: "xai:grok-4.5", provider: "xai", model: "grok-4.5" });
321+
});
312322
});
313323

314324
describe("expandMatrix", () => {

‎evals/capability/lib.ts‎

Lines changed: 15 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -527,6 +527,8 @@ export function defaultVariantId(provider?: string, model?: string): string {
527527
* - `provider/model` (slash only when no colon)
528528
* - `label=provider:model`
529529
* Empty / omitted → single default variant (caller provider/model flags).
530+
* Each expanded cell must have both provider and model (after applying
531+
* `--provider`/`--model` as cell defaults when a side is omitted).
530532
*/
531533
export function parseMatrix(
532534
matrix: string | undefined,
@@ -549,10 +551,14 @@ export function parseMatrix(
549551
if (cells.length === 0) {
550552
throw new Error("--matrix has no variants");
551553
}
552-
return cells.map((cell, index) => parseMatrixCell(cell, index));
554+
return cells.map((cell, index) => parseMatrixCell(cell, index, fallback));
553555
}
554556

555-
function parseMatrixCell(cell: string, index: number): EvalVariant {
557+
function parseMatrixCell(
558+
cell: string,
559+
index: number,
560+
fallback: { provider?: string; model?: string },
561+
): EvalVariant {
556562
let label: string | undefined;
557563
let rest = cell;
558564
const eq = cell.indexOf("=");
@@ -575,15 +581,15 @@ function parseMatrixCell(cell: string, index: number): EvalVariant {
575581
`matrix cell ${index + 1} "${cell}" must be provider:model or label=provider:model`,
576582
);
577583
}
578-
if (provider === undefined && model === undefined) {
579-
throw new Error(`matrix cell ${index + 1} "${cell}" is empty`);
584+
provider = provider ?? fallback.provider;
585+
model = model ?? fallback.model;
586+
if (provider === undefined || model === undefined) {
587+
throw new Error(
588+
`matrix cell ${index + 1} "${cell}" must specify both provider and model`,
589+
);
580590
}
581591
const id = label ?? defaultVariantId(provider, model);
582-
return {
583-
id,
584-
...(provider !== undefined ? { provider } : {}),
585-
...(model !== undefined ? { model } : {}),
586-
};
592+
return { id, provider, model };
587593
}
588594

589595
/** Cartesian product of cases × variants (cases outer for stable progress). */

‎evals/public/README.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -17,7 +17,7 @@ without vendoring a full leaderboard runner into product CI.
1717

1818
```bash
1919
# Dry plan (loads HF row, prints prompt)
20-
bun scripts/eval-public-swe-one.ts --dry-run
20+
bun scripts/eval-public-swe-one.ts --dry-run --provider <name> --model <id>
2121

2222
# Default instance: psf__requests-3362 (small repo, single failing test)
2323
bun scripts/eval-public-swe-one.ts \

‎scripts/eval-capability.test.ts‎

Lines changed: 57 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,57 @@
1+
import { describe, expect, test } from "bun:test";
2+
3+
import { parseArgs } from "./eval-capability.ts";
4+
5+
describe("parseArgs", () => {
6+
test("--help does not require provider or model", () => {
7+
const opts = parseArgs(["--help"]);
8+
expect(opts.help).toBe(true);
9+
expect(opts.provider).not.toBe("xai/thegreataxios");
10+
expect(opts.model).not.toBe("xai/thegreataxios");
11+
});
12+
13+
test("no flags throws", () => {
14+
expect(() => parseArgs([])).toThrow(/--provider/);
15+
expect(() => parseArgs([])).toThrow(/--model/);
16+
});
17+
18+
test("--provider without --model throws", () => {
19+
expect(() => parseArgs(["--provider", "foo"])).toThrow(/--model/);
20+
});
21+
22+
test("--model without --provider throws", () => {
23+
expect(() => parseArgs(["--model", "bar"])).toThrow(/--provider/);
24+
});
25+
26+
test("--provider foo --model bar parses those values", () => {
27+
const opts = parseArgs(["--provider", "foo", "--model", "bar"]);
28+
expect(opts.provider).toBe("foo");
29+
expect(opts.model).toBe("bar");
30+
});
31+
32+
test("--dry-run without pair throws", () => {
33+
expect(() => parseArgs(["--dry-run"])).toThrow(/--provider/);
34+
expect(() => parseArgs(["--dry-run"])).toThrow(/--model/);
35+
});
36+
37+
test("--matrix xai:grok-4.5 is enough without top-level flags", () => {
38+
const opts = parseArgs(["--matrix", "xai:grok-4.5"]);
39+
expect(opts.matrix).toBe("xai:grok-4.5");
40+
});
41+
42+
test("incomplete matrix cell throws", () => {
43+
expect(() => parseArgs(["--matrix", "xai:"])).toThrow(/both provider and model/);
44+
expect(() => parseArgs(["--matrix", ":grok-4.5"])).toThrow(/both provider and model/);
45+
});
46+
47+
test("parsed defaults never equal xai/thegreataxios", () => {
48+
const help = parseArgs(["--help"]);
49+
const pair = parseArgs(["--provider", "foo", "--model", "bar"]);
50+
expect(help.provider).not.toBe("xai/thegreataxios");
51+
expect(help.model).not.toBe("xai/thegreataxios");
52+
expect(pair.provider).not.toBe("xai/thegreataxios");
53+
expect(pair.model).not.toBe("xai/thegreataxios");
54+
expect(pair.provider).toBe("foo");
55+
expect(pair.model).toBe("bar");
56+
});
57+
});

‎scripts/eval-capability.ts‎

Lines changed: 40 additions & 9 deletions
Original file line numberDiff line numberDiff line change
@@ -69,6 +69,7 @@ type CliOptions = {
6969
/** Runs per case×variant cell (gate runs use 5; freeze runs use 3). */
7070
repeats: number;
7171
dryRun: boolean;
72+
help: boolean;
7273
/**
7374
* Allow a run/comparison to proceed when the resolved provider/model
7475
* differs from what was requested, instead of hard-failing.
@@ -77,12 +78,13 @@ type CliOptions = {
7778
};
7879

7980
function printUsage(): void {
80-
console.log(`Usage: bun scripts/eval-capability.ts [options]
81+
console.log(`Usage: bun scripts/eval-capability.ts --provider <name> --model <id> [options]
82+
bun scripts/eval-capability.ts --matrix <cells> [options]
8183
8284
--case <id|all> Case id (default: all)
83-
--provider <name> Provider override (single-variant run)
84-
--model <id> Model override (single-variant run)
85-
--matrix <cells> Multi-variant: "p1:m1,p2:m2" or "label=p:m,..."
85+
--provider <name> Provider name (required except --help, or --matrix with complete cells)
86+
--model <id> Model id (required except --help, or --matrix with complete cells)
87+
--matrix <cells> Multi-variant: "p1:m1,p2:m2" or "label=p:m,..."; each cell needs both sides
8688
--config <path> Settings file override
8789
--out <path> Write results JSON
8890
--baseline <path> Compare to prior results JSON
@@ -91,19 +93,20 @@ function printUsage(): void {
9193
--agent-timeout-ms <n> Wall-clock limit for runExec (default 1200000)
9294
--verify-timeout-ms <n> Wall-clock limit for verify.sh (default 120000)
9395
--repeats <n> Runs per case×variant cell (default 1; gate runs use 5)
94-
--dry-run List cases × variants only
96+
--dry-run List cases × variants only (still requires --provider/--model or --matrix)
9597
--allow-provider-fallback Allow resolved provider/model to differ from
9698
what was requested (default: hard-fail)
9799
-h, --help Show help
98100
`);
99101
}
100102

101-
function parseArgs(argv: readonly string[]): CliOptions {
103+
export function parseArgs(argv: readonly string[]): CliOptions {
102104
const opts: CliOptions = {
103105
caseSelector: "all",
104106
skipPermissions: true,
105107
repeats: 1,
106108
dryRun: false,
109+
help: false,
107110
allowProviderFallback: false,
108111
agentTimeoutMs: Number(process.env.CORBITS_EVAL_AGENT_TIMEOUT_MS ?? 1_200_000),
109112
verifyTimeoutMs: Number(process.env.CORBITS_EVAL_VERIFY_TIMEOUT_MS ?? 120_000),
@@ -118,8 +121,7 @@ function parseArgs(argv: readonly string[]): CliOptions {
118121
switch (a) {
119122
case "-h":
120123
case "--help":
121-
printUsage();
122-
process.exit(0);
124+
opts.help = true;
123125
break;
124126
case "--case":
125127
opts.caseSelector = next();
@@ -183,9 +185,32 @@ function parseArgs(argv: readonly string[]): CliOptions {
183185
throw new Error(`Unknown argument: ${a}`);
184186
}
185187
}
188+
if (!opts.help) {
189+
requireExplicitModelPair(opts);
190+
}
186191
return opts;
187192
}
188193

194+
function requireExplicitModelPair(opts: CliOptions): void {
195+
const matrix = opts.matrix?.trim();
196+
if (matrix !== undefined && matrix.length > 0) {
197+
parseMatrix(matrix, {
198+
provider: opts.provider,
199+
model: opts.model,
200+
});
201+
return;
202+
}
203+
if (!opts.provider && !opts.model) {
204+
throw new Error("missing required --provider and --model (or --matrix)");
205+
}
206+
if (!opts.provider) {
207+
throw new Error("missing required --provider");
208+
}
209+
if (!opts.model) {
210+
throw new Error("missing required --model");
211+
}
212+
}
213+
189214
function runCommand(
190215
command: string,
191216
args: readonly string[],
@@ -510,7 +535,9 @@ async function runCase(
510535

511536
const config = await loadConfig(argv, { allowUnconfigured: false });
512537
if (!config.configured) {
513-
throw new Error("Provider not configured for eval run");
538+
throw new Error(
539+
"Provider not configured. Pass --provider <name> --model <id> (or --matrix) matching a configured provider.",
540+
);
514541
}
515542

516543
const agentStarted = Date.now();
@@ -684,6 +711,10 @@ function formatMetricsLine(r: CaseResult): string {
684711

685712
async function main(): Promise<number> {
686713
const opts = parseArgs(process.argv.slice(2));
714+
if (opts.help) {
715+
printUsage();
716+
return 0;
717+
}
687718
const all = await loadEvalCases(CASES_ROOT);
688719
const selected = filterCases(all, opts.caseSelector);
689720
const variants = parseMatrix(opts.matrix, {

‎scripts/eval-public-swe-one.test.ts‎

Lines changed: 9 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -10,11 +10,16 @@ describe("parseArgs", () => {
1010
expect(opts.model).not.toBe("xai/thegreataxios");
1111
});
1212

13-
test("--dry-run does not require provider or model", () => {
14-
const opts = parseArgs(["--dry-run"]);
13+
test("--dry-run alone throws", () => {
14+
expect(() => parseArgs(["--dry-run"])).toThrow(/--provider/);
15+
expect(() => parseArgs(["--dry-run"])).toThrow(/--model/);
16+
});
17+
18+
test("--dry-run with provider and model parses", () => {
19+
const opts = parseArgs(["--dry-run", "--provider", "foo", "--model", "bar"]);
1520
expect(opts.dryRun).toBe(true);
16-
expect(opts.provider).not.toBe("xai/thegreataxios");
17-
expect(opts.model).not.toBe("xai/thegreataxios");
21+
expect(opts.provider).toBe("foo");
22+
expect(opts.model).toBe("bar");
1823
});
1924

2025
test("agent run without --provider throws", () => {
@@ -38,10 +43,7 @@ describe("parseArgs", () => {
3843

3944
test("parsed defaults never equal xai/thegreataxios", () => {
4045
const help = parseArgs(["--help"]);
41-
const dry = parseArgs(["--dry-run"]);
4246
expect(help.provider).not.toBe("xai/thegreataxios");
4347
expect(help.model).not.toBe("xai/thegreataxios");
44-
expect(dry.provider).not.toBe("xai/thegreataxios");
45-
expect(dry.model).not.toBe("xai/thegreataxios");
4648
});
4749
});

‎scripts/eval-public-swe-one.ts‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -10,7 +10,7 @@
1010
* Usage:
1111
* bun scripts/eval-public-swe-one.ts --provider <name> --model <id>
1212
* bun scripts/eval-public-swe-one.ts --instance psf__requests-3362 --provider <name> --model <id>
13-
* bun scripts/eval-public-swe-one.ts --dry-run
13+
* bun scripts/eval-public-swe-one.ts --dry-run --provider <name> --model <id>
1414
*
1515
* Optional official grading (heavy; needs Docker resources):
1616
* bun scripts/eval-public-swe-one.ts --instance … --provider <name> --model <id> --evaluate
@@ -62,8 +62,8 @@ One public SWE-bench Lite instance via Corbits product exec.
6262
6363
Options:
6464
--instance <id> SWE-bench instance_id (default: ${DEFAULT_INSTANCE})
65-
--provider <name> Provider name (required except --help / --dry-run)
66-
--model <id> Model id (required except --help / --dry-run)
65+
--provider <name> Provider name (required except --help)
66+
--model <id> Model id (required except --help)
6767
--subset <hf> HF dataset id (default: ${DEFAULT_SUBSET})
6868
--split <name> Dataset split (default: ${DEFAULT_SPLIT})
6969
--timeout-ms <n> Agent wall-clock timeout (default: ${DEFAULT_AGENT_TIMEOUT_MS})
@@ -133,7 +133,7 @@ export function parseArgs(argv: string[]): CliOptions {
133133
throw new Error(`unknown arg: ${a}`);
134134
}
135135
}
136-
if (!opts.help && !opts.dryRun) {
136+
if (!opts.help) {
137137
if (!opts.provider && !opts.model) {
138138
throw new Error("missing required --provider and --model");
139139
}

0 commit comments

Comments
 (0)