Skip to content

feat(presets): add Terminal-Bench 3 and 4 releases - #234

Merged
osolmaz merged 1 commit into
mainfrom
feat/terminal-bench-3-4-presets
Oct 8, 2026
Merged

osolmaz merged 1 commit into
mainfrom
feat/terminal-bench-3-4-presets

Conversation

@osolmaz

@osolmaz osolmaz commented Sep 25, 2026 •

Copy link
Copy Markdown
Member

Summary

The built-in catalog only offers Terminal-Bench 2.1.
This adds the official 3.0 and 4.0 releases so operators can select either benchmark without an external preset source.
Each version has a one-task diagnostic, an all-task diagnostic, and a five-attempt final preset.

What Changed

The presets use the native Harbor dataset and trial fields.
They pin the official release commits rather than a moving tag. Terminal-Bench 3.0 contains 74 tasks; 4.0 contains 66.

  • Both diagnostic one-task presets select interleaved-vigenere, which exists at both release commits.
  • The final presets use five attempts per task. Only those presets are leaderboard eligible.
  • The task-defined agent timeouts remain unchanged. A 2.1-style multiplier would extend some 4.0 tasks beyond their eight-hour limit.
  • Catalog tests check the source commit, task selection, attempts, concurrency, environment, and timeout policy.

Testing

The preset files load through the Harbor-HF catalog and validate with Harbor's pinned JobConfig schema.
No paid model inference or sandbox Job was run.

  • Verified the official release tags and counted 74 and 66 task.toml files in their pinned Git trees.
  • npm run typecheck, npm run format:check, npm run lint, npm run build, npm run check:generated, and npm audit --audit-level=low passed.
  • LANG=en_US.UTF-8 LC_ALL=en_US.UTF-8 npm test passed all 1,723 tests. The default Singapore locale changes an unrelated currency-formatting expectation in three existing web tests.
  • uv run --no-sync python scripts/check_public_privacy.py . passed using the existing local environment.

Risks

This PR is draft because the all-task presets cannot run on the current CPU-only HF Sandbox hardware. Auditing every official task.toml at both release commits found four H100 tasks in 3.0 and three in 4.0. Both releases also contain two tasks requesting 16 CPUs (CPU Upgrade has eight) and one task requesting 1,024,000 MB of storage (CPU Upgrade has 50 GB). A full official run needs a separately reviewed hardware path. Do not merge or deploy these presets as all-task support yet.

Static validation does not prove that task images build, that a model endpoint responds, or that the full run fits its cost limit. The one-task diagnostic uses a task whose Dockerfile installs system packages and runs task-supplied Python during build. Review and authorize that code before a live test. Full runs need a verified cumulative campaign limit and an approved inference connection for each model.

Co-Authored-By: GPT-6 Sol <noreply@pi.dev>
Generated-By: pi 0.87.1 (https://pi.dev)
@osolmaz
osolmaz marked this pull request as draft September 25, 2026 12:53
@osolmaz
osolmaz marked this pull request as ready for review October 8, 2026 06:09
@osolmaz
osolmaz merged commit 08aab7c into main Oct 8, 2026
1 check passed
@osolmaz
osolmaz deleted the feat/terminal-bench-3-4-presets branch October 8, 2026 06:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant