(2/n) Organize Benchmark Execution Into Responsibility-Owned Packages - #80
Open
kargibora wants to merge 3 commits into
Open
(2/n) Organize Benchmark Execution Into Responsibility-Owned Packages#80kargibora wants to merge 3 commits into
kargibora wants to merge 3 commits into
Conversation
geoalgo
reviewed
Aug 3, 2026
geoalgo
left a comment
Collaborator
There was a problem hiding this comment.
LGTM, I have only minor comments.
Comment on lines
+18
to
+21
| def slugify(value: str) -> str: | ||
| """Return a filesystem-safe model, task, or benchmark name.""" | ||
| slug = re.sub(r"[^a-z0-9]+", "-", value.lower()).strip("-") | ||
| return slug or "value" |
Collaborator
There was a problem hiding this comment.
not sure about the name (I only vaguely could guess from the name).
Given that we will have it at a bunch of place perhaps we can choose a more telling name like safe_string? or safe_filename?
Collaborator
Author
There was a problem hiding this comment.
I agree. Slugify is a technical term for generating a safe name from a string with spaces however we dont have to set it up like that. safe_filename sounds better to me
Collaborator
Author
There was a problem hiding this comment.
61e23f2 renames the function to safe_filename
kargibora
marked this pull request as ready for review
August 4, 2026 09:33
kargibora
force-pushed
the
refactor/benchmark-packages-main
branch
from
August 4, 2026 11:16
61e23f2 to
4cefd83
Compare
Move existing implementations into responsibility-owned packages without changing benchmark behavior.
Route generate-and-evaluate tasks through a dedicated registry and thin shared runner.
Clearer name for the filesystem-safe string helper used across benchmark runners, per review.
geoalgo
force-pushed
the
refactor/benchmark-packages-main
branch
from
August 5, 2026 13:32
4cefd83 to
ab0e602
Compare
geoalgo
approved these changes
Aug 5, 2026
geoalgo
left a comment
Collaborator
There was a problem hiding this comment.
LGTM, we just need to fix ruff / CI before merging, feel free to merge once this is done.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The previous PR introduced a shared benchmark runner interface, but benchmark implementations and related utilities are still distributed across the root of
judgearena.For example, pairwise evaluation, ELO execution, MT-Bench, datasets, and run metadata currently live in unrelated root-
level modules. This makes ownership unclear and increases the number of unrelated files that must be changed when adding or modifying a benchmark.
This PR reorganizes the existing implementation into responsibility-owned packages and moves benchmark dispatch out of the pairwise runner.
Take a look at SoP to understand the reason and shortcomings of the current approach.
What changed
Benchmark implementations are now grouped under:
Important changes include:
Most of the diff consists of file moves and import updates. Existing benchmark protocols, prompts, configuration, and
evaluation behavior are preserved.
Why this helps
A benchmark can now own its runner and utilities without adding more logic to the pairwise implementation or the package
root. It also gives datasets and artifacts stable locations that can be reused by future benchmark integrations.
Notes