Repository navigation
feat(presets): add Terminal-Bench 3 and 4 releases - #234
Merged
Merged
Conversation
Co-Authored-By: GPT-6 Sol <noreply@pi.dev> Generated-By: pi 0.87.1 (https://pi.dev)
osolmaz
marked this pull request as draft
September 25, 2026 12:53
osolmaz
marked this pull request as ready for review
October 8, 2026 06:09
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The built-in catalog only offers Terminal-Bench 2.1.
This adds the official 3.0 and 4.0 releases so operators can select either benchmark without an external preset source.
Each version has a one-task diagnostic, an all-task diagnostic, and a five-attempt final preset.
What Changed
The presets use the native Harbor dataset and trial fields.
They pin the official release commits rather than a moving tag. Terminal-Bench 3.0 contains 74 tasks; 4.0 contains 66.
interleaved-vigenere, which exists at both release commits.Testing
The preset files load through the Harbor-HF catalog and validate with Harbor's pinned
JobConfigschema.No paid model inference or sandbox Job was run.
task.tomlfiles in their pinned Git trees.npm run typecheck,npm run format:check,npm run lint,npm run build,npm run check:generated, andnpm audit --audit-level=lowpassed.LANG=en_US.UTF-8 LC_ALL=en_US.UTF-8 npm testpassed all 1,723 tests. The default Singapore locale changes an unrelated currency-formatting expectation in three existing web tests.uv run --no-sync python scripts/check_public_privacy.py .passed using the existing local environment.Risks
This PR is draft because the all-task presets cannot run on the current CPU-only HF Sandbox hardware. Auditing every official
task.tomlat both release commits found four H100 tasks in 3.0 and three in 4.0. Both releases also contain two tasks requesting 16 CPUs (CPU Upgrade has eight) and one task requesting 1,024,000 MB of storage (CPU Upgrade has 50 GB). A full official run needs a separately reviewed hardware path. Do not merge or deploy these presets as all-task support yet.Static validation does not prove that task images build, that a model endpoint responds, or that the full run fits its cost limit. The one-task diagnostic uses a task whose Dockerfile installs system packages and runs task-supplied Python during build. Review and authorize that code before a live test. Full runs need a verified cumulative campaign limit and an approved inference connection for each model.