Skip to content

Zero-activity rollouts are recorded as graded 'timeout' (reward 0) with a fabricated elapsed time #1071

Description

@yuchenlwu

Observed

Running FrontierPhysics tasks (docker sandbox, claude-agent-acp), some rollouts die with zero tool calls, zero trajectory events after only 2-70 minutes of wall time, yet are recorded as:

  • error_category: "timeout" (the graded category — verifier runs on an untouched workspace and scores 0), and
  • error: "Agent timed out after 14400s" where 14400s is the task's configured agent.timeout_sec — the actual elapsed time was 303s in our specimen (timing: {environment_setup: 166.2, verifier: 3.3, total: 303.5}, n_tool_calls: 0, trajectory_source: null).

We hit 10 such rollouts across a 180-rollout batch (both claude-fable-5 and claude-opus-5).

Why it matters

  1. A rollout in which the agent never produced a single event is an infra failure, not a model result — scoring it 0 systematically deflates model scores. In our batch it moved the macro-average by ~2x (0.075 → 0.154 after exclusion).
  2. The error string reports the configured budget as if it were measured ("timed out after 14400s" on a 5-minute rollout), which makes triage actively misleading.

Suggestion

  • Classify a prompt-timeout with zero agent activity (no tool calls, no chunks, no trajectory) as an infra/retryable category rather than the graded timeout.
  • Report measured elapsed time in the timeout message (or both configured and measured).

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions