You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This is a BenchGuard integration/evidence-contract tracking issue, not one atomic runtime bug. It tracks two validator regressions and one executable state-model contract gap found while I was prototyping #1050 and #1051 with an agent:
artifacts cannot prove that grading ran in a declared separated verifier;
after positive agent-process quiescence evidence is recorded, an interrupted tool can remain permanently pending with no framework-owned terminal lifecycle evidence; and
the experiment-review validator rejects a native ACP timeout that is a candidate for completeness under the proposed BenchGuard contract.
These are related by BenchGuard's evidence requirements, but they should not be implemented as one large patch. Each section below has reproducible evidence—two failing validator regressions and one executable state-model demonstration—and its own acceptance boundary. The exact schema and status vocabulary remain design choices.
The prototypes are useful design evidence. I do not think either PR should be merged as-is: #1050's verifier predicate accepts setup-only execs, while #1051 combines transport cancellation, process hardening, lifecycle synthesis, verifier behavior, and artifact validation.
Dependencies
Two standalone BenchFlow bugs should remain separate issues:
This tracking issue must not silently choose behavior for either dependency.
Stable reproduction baseline
The commands below use 656ec91, the common base of #1050 and #1051, so results do not depend on later prototype changes:
git clone https://github.com/benchflow-ai/benchflow.git
cd benchflow
git switch --detach 656ec915574fb981becf00599e8acc323732c3aa
uv sync --extra dev --locked
Later unrelated mainline commits do not change the referenced validator/runtime paths at the time of writing.
Run Gap 1 and Gap 3 from separate clean checkouts. Both reproductions patch tests/test_benchflow_experiment_review_validator.py; changes left by one reproduction would otherwise interfere with switching to the other pinned baseline.
Gap 1: artifacts cannot prove separated-verifier grading ran
Mainline validate_rollout() validates generic rollout artifacts but does not reconcile that declaration with runtime evidence that the grading attempt ran in benchguard-verifier.
Apply #1050's prototype test additions as a test-only diff, leaving the base validator unchanged. The diff adds supporting fixtures and tests beyond the single target selected below:
#1050 accepts the existence of a successful sidecar-start action plus any SandboxExec with phase: "Verify", service: "benchguard-verifier", and transport status: "ok":
That stream contains verifier setup only; no grading attempt. It nevertheless satisfies #1050's validate_separate_verifier_actions(). status: "ok" also describes the sandbox exec API call, while process outcome is recorded separately as return_code.
BenchGuard impact
A rollout can be certified as isolated without evidence that grading crossed the intended service boundary.
A crash after sidecar startup or setup can be indistinguishable from completed grading.
Experiment review can rely on an isolation claim stronger than the artifact bundle supports.
Acceptance boundary
Explicit separate mode requires grading-attempt evidence tied to the declared verifier service.
Sidecar startup alone and setup-only execs are rejected.
Attempt identity, ordering, terminal outcome, and process result can be reconciled.
Missing, malformed, contradictory, or incomplete relevant evidence fails closed.
Undeclared/default compatibility is documented and tested separately.
Possible direction: emit typed grading-attempt evidence at the runtime point that owns verifier execution. This is one option, not a prescribed event schema.
Gap 2: timeout tools lack terminal lifecycle evidence after a positive quiescence observation
Problem
At timeout, mainline snapshots pending tool IDs and sets terminal_trajectory_complete = false. Even if later verifier hardening records a positive post-kill quiescence observation under the adopted policy, the already-published trajectory still contains pending/in_progress; nothing records the narrower framework conclusion that the interrupted execution has ended while past effects remain unknown.
Rollout.verify() publishes the trajectory before harden_before_verify() runs.
Mainline's process cleanup is in _kill_sandbox_user_procs(), but the standalone quiescence issue must first define and enforce the required checked observation.
Executable state-model demonstration
This demonstrates the missing state transition deterministically. It is a contract demonstration, not an end-to-end quiescence reproduction: agent_quiescence_observed is an explicit assumed external premise governed by the standalone quiescence issue.
Run from the stable baseline checkout:
uv run python - <<'PY'from benchflow.acp.session import ACPSessionfrom benchflow.trajectories._capture import _capture_session_trajectorysession = ACPSession("timeout-repro")session.handle_update({ "sessionUpdate": "tool_call", "toolCallId": "tc-1", "title": "long command", "kind": "bash",})session.record_agent_timeout( timeout_sec=45.0, pending_tool_call_ids=["tc-1"], terminal_trajectory_complete=False,)# External premise supplied by the adopted checked-quiescence policy.agent_quiescence_observed = Truerows = _capture_session_trajectory(session)print(rows)assert agent_quiescence_observedassert rows[0]["status"] not in {"pending", "in_progress"}assert rows[1]["terminal_trajectory_complete"] is TruePY
Actual: assertions fail. Tool remains pending; timeout remains incomplete. Mainline has no state transition that can consume positive quiescence evidence.
Expected under the adopted policy: after genuine adapter updates have been drained and a checked observation supplies positive quiescence evidence, artifacts distinguish “framework ended this execution; effects unknown” from a live pending tool. They must not claim that the command never had effects.
Prototype reference: #1051 introduced one possible implementation in ddab74c, with focused fixture TestTimeoutTerminalization. Treat status: "cancelled", effects: "unknown", and terminal_source: "benchflow" as prototype vocabulary, not an accepted schema decision.
BenchGuard impact
Timeouts with positive quiescence evidence remain structurally partial forever, so legitimate timeout trials cannot become countable under a terminal evidence contract.
Dropping every such trial biases experiments toward easy/cooperative completions.
Treating pending as terminal would instead overstate evidence and erase the distinction between adapter-reported outcome and framework-owned lifecycle.
Acceptance boundary
Genuine adapter updates preserved by the ACP cancellation issue take precedence.
Framework-owned evidence is added only after the checked observation required by the quiescence issue supplies positive quiescence evidence.
Evidence says execution cannot continue; it does not claim known command effects or an adapter-reported outcome.
Provenance distinguishes adapter evidence from BenchFlow inference.
Wall-clock and idle timeout paths use equivalent semantics.
If the required observation is unavailable or reports surviving relevant processes, the tool remains unresolved and the trajectory remains partial.
Finalization is idempotent and visible to verifier plus exported artifacts.
Possible direction: add a framework-owned terminal lifecycle state or annotation with explicit unknown effects. A new status, an orthogonal lifecycle field, or another representation may be preferable to the prototype's synthetic cancelled status.
Gap 3: validator rejects candidate-complete native ACP timeouts
Problem
The experiment-review validator has an explicit exception for native subscription rollouts without provider llm_trajectory.jsonl, but it assumes every accepted row is a successful agent completion. It requires:
A real timeout must report the opposite completion/error shape. Even when it has positive native usage, timing, verifier reward, explicit timeout metadata, terminal ACP trajectory evidence, and training_ready: false, the validator rejects a candidate complete under the proposed BenchGuard contract.
Deterministic reproduction
In a separate clean checkout, start from #1050 immediately before its timeout-validator commit. Apply the commit's prototype test additions as a test-only diff; the diff adds supporting fixtures and tests beyond the single target selected below:
Actual: test fails. Validator reports that the native subscription row carries an error, is not completed, and has an unexpected stop condition.
Expected under the proposed BenchGuard contract: a timeout satisfying the candidate completeness requirements is accepted as countable experiment evidence while remaining explicitly non-training-ready and non-successful.
Positive usage, timing, verifier reward/score, timeout diagnostic, and terminal trajectory evidence are required and mutually consistent.
Accepted timeout remains training_ready: false; no prompt/completion/training trajectory is synthesized.
Pending/in-progress tools or contradictory completeness claims are rejected.
Infra/provider failures are not reclassified as normal agent timeouts.
Wall-clock and idle timeout representation is handled deliberately rather than accidentally accepting only one diagnostic shape.
Possible direction: validate an explicit timeout artifact variant instead of weakening successful native-subscription requirements. The predicate added by 181ce4e is a prototype, not necessarily the best final contract.
Tracking completion
This issue is complete when all three gaps have regression coverage and documented artifact semantics, after the standalone ACP-cancellation and process-quiescence dependencies land where required. Separate implementation PRs are expected.
Suggested checklist:
Gap 1: typed, correlated evidence proves separated grading execution.
Gap 2: interrupted tools receive honest framework-owned lifecycle evidence only after positive quiescence evidence under the adopted policy.
Gap 3: native ACP timeouts satisfying the adopted completeness contract are countable but never training-ready.
Cross-gap fixtures reject missing, malformed, pending, and contradictory evidence.
Docs state which claims come from adapter events, BenchFlow runtime observations, and BenchGuard validation.
Non-goals
Choosing the ACP cancellation transport/grace implementation; track in the standalone timeout bug.
Defining the checked quiescence policy or treating pkill success alone as sufficient evidence; track in the standalone quiescence bug.
Treating every timeout as complete.
Inferring tool effects from process death.
Parsing verifier stdout to infer where grading ran.
Closing #1050 and #1051 in favor of this tracking issue plus the two standalone bug issues would preserve their discussion and attribution without presenting either prototype branch as a merge-ready implementation.
Summary
This is a BenchGuard integration/evidence-contract tracking issue, not one atomic runtime bug. It tracks two validator regressions and one executable state-model contract gap found while I was prototyping #1050 and #1051 with an agent:
pendingwith no framework-owned terminal lifecycle evidence; andThese are related by BenchGuard's evidence requirements, but they should not be implemented as one large patch. Each section below has reproducible evidence—two failing validator regressions and one executable state-model demonstration—and its own acceptance boundary. The exact schema and status vocabulary remain design choices.
The prototypes are useful design evidence. I do not think either PR should be merged as-is: #1050's verifier predicate accepts setup-only execs, while #1051 combines transport cancellation, process hardening, lifecycle synthesis, verifier behavior, and artifact validation.
Dependencies
Two standalone BenchFlow bugs should remain separate issues:
pkillAPI call is not sufficient.This tracking issue must not silently choose behavior for either dependency.
Stable reproduction baseline
The commands below use
656ec91, the common base of #1050 and #1051, so results do not depend on later prototype changes:git clone https://github.com/benchflow-ai/benchflow.git cd benchflow git switch --detach 656ec915574fb981becf00599e8acc323732c3aa uv sync --extra dev --lockedLater unrelated mainline commits do not change the referenced validator/runtime paths at the time of writing.
Run Gap 1 and Gap 3 from separate clean checkouts. Both reproductions patch
tests/test_benchflow_experiment_review_validator.py; changes left by one reproduction would otherwise interfere with switching to the other pinned baseline.Gap 1: artifacts cannot prove separated-verifier grading ran
Problem
A BenchGuard preflight may explicitly record:
{ "report": { "task_binding": { "facts": { "verifier_environment_mode": "separate" } } } }Mainline
validate_rollout()validates generic rollout artifacts but does not reconcile that declaration with runtime evidence that the grading attempt ran inbenchguard-verifier.Code reference:
validate_rollout()at the common base. There is no separated-verifier evidence check in this path.Deterministic reproduction
Apply #1050's prototype test additions as a test-only diff, leaving the base validator unchanged. The diff adds supporting fixtures and tests beyond the single target selected below:
git diff 656ec915574fb981becf00599e8acc323732c3aa \ 27b30b76ea97391287a9e4f8d4f5279803f1ee79 \ -- tests/test_benchflow_experiment_review_validator.py | git apply uv run --extra dev python -m pytest -q \ tests/test_benchflow_experiment_review_validator.py::test_validator_rejects_separate_verifier_without_sidecar_actionsActual: test fails because report remains healthy even though preflight declares separate verification and
benchguard/action_records.jsonlis absent.Expected: validator reports missing or incomplete proof for the declared verifier mode.
Fixture source:
test_validator_rejects_separate_verifier_without_sidecar_actions.Why #1050's predicate is still a false positive
#1050 accepts the existence of a successful sidecar-start action plus any
SandboxExecwithphase: "Verify",service: "benchguard-verifier", and transportstatus: "ok":{"event_type":"StartBenchGuardVerifierService","phase":"Verify","service":"benchguard-verifier","status":"ok"} {"event_type":"SandboxExec","phase":"Verify","service":"benchguard-verifier","status":"ok","command_preview":"mkdir -p /logs/verifier && chmod 777 /logs/verifier","return_code":0}That stream contains verifier setup only; no grading attempt. It nevertheless satisfies #1050's
validate_separate_verifier_actions().status: "ok"also describes the sandbox exec API call, while process outcome is recorded separately asreturn_code.BenchGuard impact
Acceptance boundary
Possible direction: emit typed grading-attempt evidence at the runtime point that owns verifier execution. This is one option, not a prescribed event schema.
Gap 2: timeout tools lack terminal lifecycle evidence after a positive quiescence observation
Problem
At timeout, mainline snapshots pending tool IDs and sets
terminal_trajectory_complete = false. Even if later verifier hardening records a positive post-kill quiescence observation under the adopted policy, the already-published trajectory still containspending/in_progress; nothing records the narrower framework conclusion that the interrupted execution has ended while past effects remain unknown.Relevant ordering and state:
_agent_prompt_timeout_error()marks the timeout incomplete when pending tools exist.ACPSession.record_agent_timeout()records no later process-lifecycle transition.Rollout.verify()publishes the trajectory beforeharden_before_verify()runs._kill_sandbox_user_procs(), but the standalone quiescence issue must first define and enforce the required checked observation.Executable state-model demonstration
This demonstrates the missing state transition deterministically. It is a contract demonstration, not an end-to-end quiescence reproduction:
agent_quiescence_observedis an explicit assumed external premise governed by the standalone quiescence issue.Run from the stable baseline checkout:
Actual: assertions fail. Tool remains
pending; timeout remains incomplete. Mainline has no state transition that can consume positive quiescence evidence.Expected under the adopted policy: after genuine adapter updates have been drained and a checked observation supplies positive quiescence evidence, artifacts distinguish “framework ended this execution; effects unknown” from a live pending tool. They must not claim that the command never had effects.
Prototype reference: #1051 introduced one possible implementation in
ddab74c, with focused fixtureTestTimeoutTerminalization. Treatstatus: "cancelled",effects: "unknown", andterminal_source: "benchflow"as prototype vocabulary, not an accepted schema decision.BenchGuard impact
pendingas terminal would instead overstate evidence and erase the distinction between adapter-reported outcome and framework-owned lifecycle.Acceptance boundary
Possible direction: add a framework-owned terminal lifecycle state or annotation with explicit unknown effects. A new status, an orthogonal lifecycle field, or another representation may be preferable to the prototype's synthetic
cancelledstatus.Gap 3: validator rejects candidate-complete native ACP timeouts
Problem
The experiment-review validator has an explicit exception for native subscription rollouts without provider
llm_trajectory.jsonl, but it assumes every accepted row is a successful agent completion. It requires:error is None;is_completed is True; andstop_condition == "agent_completed".Code reference:
validate_results_row()native-subscription branch.A real timeout must report the opposite completion/error shape. Even when it has positive native usage, timing, verifier reward, explicit timeout metadata, terminal ACP trajectory evidence, and
training_ready: false, the validator rejects a candidate complete under the proposed BenchGuard contract.Deterministic reproduction
In a separate clean checkout, start from #1050 immediately before its timeout-validator commit. Apply the commit's prototype test additions as a test-only diff; the diff adds supporting fixtures and tests beyond the single target selected below:
git switch --detach 158d7224c7b7d7dc7c4951bef93bb7aed230c121 git diff 158d7224c7b7d7dc7c4951bef93bb7aed230c121 \ 181ce4e7b4987ec003caf36e854463f95f319bbf \ -- tests/test_benchflow_experiment_review_validator.py | git apply uv run --extra dev python -m pytest -q \ tests/test_benchflow_experiment_review_validator.py::test_validator_accepts_complete_native_subscription_timeoutActual: test fails. Validator reports that the native subscription row carries an error, is not completed, and has an unexpected stop condition.
Expected under the proposed BenchGuard contract: a timeout satisfying the candidate completeness requirements is accepted as countable experiment evidence while remaining explicitly non-training-ready and non-successful.
Fixture source:
_native_subscription_timeout_rollout()and acceptance test.BenchGuard impact
agent_completedmerely to satisfy a success-only validator shape.Acceptance boundary
is_completed: false, and timeout-compatible stop condition.training_ready: false; no prompt/completion/training trajectory is synthesized.Possible direction: validate an explicit timeout artifact variant instead of weakening successful native-subscription requirements. The predicate added by
181ce4eis a prototype, not necessarily the best final contract.Tracking completion
This issue is complete when all three gaps have regression coverage and documented artifact semantics, after the standalone ACP-cancellation and process-quiescence dependencies land where required. Separate implementation PRs are expected.
Suggested checklist:
Non-goals
pkillsuccess alone as sufficient evidence; track in the standalone quiescence bug.Provenance
fix: validate separate verifier evidence, opened by me (kywch) from agent-assisted BenchGuard work27b30b7181ce4ePreserve native ACP artifacts at prompt timeout, opened by me (kywch) from the same agent-assisted work0b5ccb1ddab74cClosing #1050 and #1051 in favor of this tracking issue plus the two standalone bug issues would preserve their discussion and attribution without presenting either prototype branch as a merge-ready implementation.