Skip to content

[Tracking] Define three BenchGuard artifact-evidence contracts #1079

Description

@kywch

Summary

This is a BenchGuard integration/evidence-contract tracking issue, not one atomic runtime bug. It tracks two validator regressions and one executable state-model contract gap found while I was prototyping #1050 and #1051 with an agent:

  1. artifacts cannot prove that grading ran in a declared separated verifier;
  2. after positive agent-process quiescence evidence is recorded, an interrupted tool can remain permanently pending with no framework-owned terminal lifecycle evidence; and
  3. the experiment-review validator rejects a native ACP timeout that is a candidate for completeness under the proposed BenchGuard contract.

These are related by BenchGuard's evidence requirements, but they should not be implemented as one large patch. Each section below has reproducible evidence—two failing validator regressions and one executable state-model demonstration—and its own acceptance boundary. The exact schema and status vocabulary remain design choices.

The prototypes are useful design evidence. I do not think either PR should be merged as-is: #1050's verifier predicate accepts setup-only execs, while #1051 combines transport cancellation, process hardening, lifecycle synthesis, verifier behavior, and artifact validation.

Dependencies

Two standalone BenchFlow bugs should remain separate issues:

This tracking issue must not silently choose behavior for either dependency.

Stable reproduction baseline

The commands below use 656ec91, the common base of #1050 and #1051, so results do not depend on later prototype changes:

git clone https://github.com/benchflow-ai/benchflow.git
cd benchflow
git switch --detach 656ec915574fb981becf00599e8acc323732c3aa
uv sync --extra dev --locked

Later unrelated mainline commits do not change the referenced validator/runtime paths at the time of writing.

Run Gap 1 and Gap 3 from separate clean checkouts. Both reproductions patch tests/test_benchflow_experiment_review_validator.py; changes left by one reproduction would otherwise interfere with switching to the other pinned baseline.


Gap 1: artifacts cannot prove separated-verifier grading ran

Problem

A BenchGuard preflight may explicitly record:

{
  "report": {
    "task_binding": {
      "facts": {
        "verifier_environment_mode": "separate"
      }
    }
  }
}

Mainline validate_rollout() validates generic rollout artifacts but does not reconcile that declaration with runtime evidence that the grading attempt ran in benchguard-verifier.

Code reference: validate_rollout() at the common base. There is no separated-verifier evidence check in this path.

Deterministic reproduction

Apply #1050's prototype test additions as a test-only diff, leaving the base validator unchanged. The diff adds supporting fixtures and tests beyond the single target selected below:

git diff 656ec915574fb981becf00599e8acc323732c3aa \
  27b30b76ea97391287a9e4f8d4f5279803f1ee79 \
  -- tests/test_benchflow_experiment_review_validator.py | git apply

uv run --extra dev python -m pytest -q \
  tests/test_benchflow_experiment_review_validator.py::test_validator_rejects_separate_verifier_without_sidecar_actions

Actual: test fails because report remains healthy even though preflight declares separate verification and benchguard/action_records.jsonl is absent.

Expected: validator reports missing or incomplete proof for the declared verifier mode.

Fixture source: test_validator_rejects_separate_verifier_without_sidecar_actions.

Why #1050's predicate is still a false positive

#1050 accepts the existence of a successful sidecar-start action plus any SandboxExec with phase: "Verify", service: "benchguard-verifier", and transport status: "ok":

{"event_type":"StartBenchGuardVerifierService","phase":"Verify","service":"benchguard-verifier","status":"ok"}
{"event_type":"SandboxExec","phase":"Verify","service":"benchguard-verifier","status":"ok","command_preview":"mkdir -p /logs/verifier && chmod 777 /logs/verifier","return_code":0}

That stream contains verifier setup only; no grading attempt. It nevertheless satisfies #1050's validate_separate_verifier_actions(). status: "ok" also describes the sandbox exec API call, while process outcome is recorded separately as return_code.

BenchGuard impact

  • A rollout can be certified as isolated without evidence that grading crossed the intended service boundary.
  • A crash after sidecar startup or setup can be indistinguishable from completed grading.
  • Experiment review can rely on an isolation claim stronger than the artifact bundle supports.

Acceptance boundary

  • Explicit separate mode requires grading-attempt evidence tied to the declared verifier service.
  • Sidecar startup alone and setup-only execs are rejected.
  • Attempt identity, ordering, terminal outcome, and process result can be reconciled.
  • Missing, malformed, contradictory, or incomplete relevant evidence fails closed.
  • Existing reward semantics, including valid reward-producing nonzero verifier exits, remain representable.
  • Undeclared/default compatibility is documented and tested separately.

Possible direction: emit typed grading-attempt evidence at the runtime point that owns verifier execution. This is one option, not a prescribed event schema.


Gap 2: timeout tools lack terminal lifecycle evidence after a positive quiescence observation

Problem

At timeout, mainline snapshots pending tool IDs and sets terminal_trajectory_complete = false. Even if later verifier hardening records a positive post-kill quiescence observation under the adopted policy, the already-published trajectory still contains pending/in_progress; nothing records the narrower framework conclusion that the interrupted execution has ended while past effects remain unknown.

Relevant ordering and state:

Executable state-model demonstration

This demonstrates the missing state transition deterministically. It is a contract demonstration, not an end-to-end quiescence reproduction: agent_quiescence_observed is an explicit assumed external premise governed by the standalone quiescence issue.

Run from the stable baseline checkout:

uv run python - <<'PY'
from benchflow.acp.session import ACPSession
from benchflow.trajectories._capture import _capture_session_trajectory

session = ACPSession("timeout-repro")
session.handle_update({
    "sessionUpdate": "tool_call",
    "toolCallId": "tc-1",
    "title": "long command",
    "kind": "bash",
})
session.record_agent_timeout(
    timeout_sec=45.0,
    pending_tool_call_ids=["tc-1"],
    terminal_trajectory_complete=False,
)

# External premise supplied by the adopted checked-quiescence policy.
agent_quiescence_observed = True
rows = _capture_session_trajectory(session)
print(rows)

assert agent_quiescence_observed
assert rows[0]["status"] not in {"pending", "in_progress"}
assert rows[1]["terminal_trajectory_complete"] is True
PY

Actual: assertions fail. Tool remains pending; timeout remains incomplete. Mainline has no state transition that can consume positive quiescence evidence.

Expected under the adopted policy: after genuine adapter updates have been drained and a checked observation supplies positive quiescence evidence, artifacts distinguish “framework ended this execution; effects unknown” from a live pending tool. They must not claim that the command never had effects.

Prototype reference: #1051 introduced one possible implementation in ddab74c, with focused fixture TestTimeoutTerminalization. Treat status: "cancelled", effects: "unknown", and terminal_source: "benchflow" as prototype vocabulary, not an accepted schema decision.

BenchGuard impact

  • Timeouts with positive quiescence evidence remain structurally partial forever, so legitimate timeout trials cannot become countable under a terminal evidence contract.
  • Dropping every such trial biases experiments toward easy/cooperative completions.
  • Treating pending as terminal would instead overstate evidence and erase the distinction between adapter-reported outcome and framework-owned lifecycle.

Acceptance boundary

  • Genuine adapter updates preserved by the ACP cancellation issue take precedence.
  • Framework-owned evidence is added only after the checked observation required by the quiescence issue supplies positive quiescence evidence.
  • Evidence says execution cannot continue; it does not claim known command effects or an adapter-reported outcome.
  • Provenance distinguishes adapter evidence from BenchFlow inference.
  • Wall-clock and idle timeout paths use equivalent semantics.
  • If the required observation is unavailable or reports surviving relevant processes, the tool remains unresolved and the trajectory remains partial.
  • Finalization is idempotent and visible to verifier plus exported artifacts.

Possible direction: add a framework-owned terminal lifecycle state or annotation with explicit unknown effects. A new status, an orthogonal lifecycle field, or another representation may be preferable to the prototype's synthetic cancelled status.


Gap 3: validator rejects candidate-complete native ACP timeouts

Problem

The experiment-review validator has an explicit exception for native subscription rollouts without provider llm_trajectory.jsonl, but it assumes every accepted row is a successful agent completion. It requires:

  • error is None;
  • is_completed is True; and
  • stop_condition == "agent_completed".

Code reference: validate_results_row() native-subscription branch.

A real timeout must report the opposite completion/error shape. Even when it has positive native usage, timing, verifier reward, explicit timeout metadata, terminal ACP trajectory evidence, and training_ready: false, the validator rejects a candidate complete under the proposed BenchGuard contract.

Deterministic reproduction

In a separate clean checkout, start from #1050 immediately before its timeout-validator commit. Apply the commit's prototype test additions as a test-only diff; the diff adds supporting fixtures and tests beyond the single target selected below:

git switch --detach 158d7224c7b7d7dc7c4951bef93bb7aed230c121
git diff 158d7224c7b7d7dc7c4951bef93bb7aed230c121 \
  181ce4e7b4987ec003caf36e854463f95f319bbf \
  -- tests/test_benchflow_experiment_review_validator.py | git apply

uv run --extra dev python -m pytest -q \
  tests/test_benchflow_experiment_review_validator.py::test_validator_accepts_complete_native_subscription_timeout

Actual: test fails. Validator reports that the native subscription row carries an error, is not completed, and has an unexpected stop condition.

Expected under the proposed BenchGuard contract: a timeout satisfying the candidate completeness requirements is accepted as countable experiment evidence while remaining explicitly non-training-ready and non-successful.

Fixture source: _native_subscription_timeout_rollout() and acceptance test.

BenchGuard impact

  • Native ACP timeout cells cannot pass artifact review even when all available evidence is complete and internally consistent.
  • Excluding timeout outcomes creates completion/easy-task selection bias.
  • Operators may be tempted to mislabel a timeout as agent_completed merely to satisfy a success-only validator shape.

Acceptance boundary

  • Validator distinguishes successful completion from a timeout that satisfies the adopted BenchGuard completeness contract.
  • Accepted timeout retains non-null timeout error, is_completed: false, and timeout-compatible stop condition.
  • Positive usage, timing, verifier reward/score, timeout diagnostic, and terminal trajectory evidence are required and mutually consistent.
  • Accepted timeout remains training_ready: false; no prompt/completion/training trajectory is synthesized.
  • Pending/in-progress tools or contradictory completeness claims are rejected.
  • Infra/provider failures are not reclassified as normal agent timeouts.
  • Wall-clock and idle timeout representation is handled deliberately rather than accidentally accepting only one diagnostic shape.

Possible direction: validate an explicit timeout artifact variant instead of weakening successful native-subscription requirements. The predicate added by 181ce4e is a prototype, not necessarily the best final contract.


Tracking completion

This issue is complete when all three gaps have regression coverage and documented artifact semantics, after the standalone ACP-cancellation and process-quiescence dependencies land where required. Separate implementation PRs are expected.

Suggested checklist:

  • Gap 1: typed, correlated evidence proves separated grading execution.
  • Gap 2: interrupted tools receive honest framework-owned lifecycle evidence only after positive quiescence evidence under the adopted policy.
  • Gap 3: native ACP timeouts satisfying the adopted completeness contract are countable but never training-ready.
  • Cross-gap fixtures reject missing, malformed, pending, and contradictory evidence.
  • Docs state which claims come from adapter events, BenchFlow runtime observations, and BenchGuard validation.

Non-goals

  • Choosing the ACP cancellation transport/grace implementation; track in the standalone timeout bug.
  • Defining the checked quiescence policy or treating pkill success alone as sufficient evidence; track in the standalone quiescence bug.
  • Treating every timeout as complete.
  • Inferring tool effects from process death.
  • Parsing verifier stdout to infer where grading ran.
  • Making native timeout artifacts training-ready.
  • Merging all changes from fix: validate separate verifier evidence #1050 or Preserve native ACP artifacts at prompt timeout #1051 into one replacement PR.

Provenance

Closing #1050 and #1051 in favor of this tracking issue plus the two standalone bug issues would preserve their discussion and attribution without presenting either prototype branch as a merge-ready implementation.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions