Skip to content

watchdog restarts a run that finished on quiescence five times, leaving five empty run directories; --wait returns while the worker is still running #6089

Description

@chernistry

Problem

A run that finishes cleanly is restarted five times by the watchdog, and each restart opens a new run directory with nothing in it. Seen on the released build, one goal, one worker:

uvx --from "bernstein[trace]==3.20.0" bernstein run --port 8077 --worker backend --model claude-haiku-4-5-20251001 --auto-approve --quiet --wait

.sdd/runtime/spawner.log and .sdd/runtime/watchdog.log, same run, times local:

20:01:03  Agent backend-99365805 finished task 7f8f9bdabf7a, process reaped
20:01:05  Quiescence confirmed after 2.0s settle window (tick #7, open=0 agents=0, no active holds) - self-stopping
20:01:09  Orchestrator stopped (replay: .sdd/runs/20260920-165918/journal.jsonl, ...)
20:01:12  WARNING Orchestrator (PID 55681) is dead, restarting...          <- watchdog.log
20:01:14  Orchestrator started (poll=3s, max_agents=1, server=http://127.0.0.1:8077)
20:01:14  WAL recovery: found 2 uncommitted entries from previous run(s)
20:01:16  Quiescence confirmed after 2.0s settle window (tick #1, open=0 agents=0, no active holds) - self-stopping
20:01:19  Orchestrator stopped (replay: .sdd/runs/20260920-170114/journal.jsonl, ...)
20:01:22  WARNING Orchestrator (PID 59343) is dead, restarting...
   ... same at 20:01:32, 20:01:42, 20:01:52 ...

Result on disk: .sdd/runs/ holds the real run (20260920-165918, 18 events, run_completed) plus five directories 20260920-1701{14,24,34,44,54} whose journals are run_started, loaded_extension_set, tick_start, task_completed, run_quiescence, run_completed and nothing else. The loop only ends because max_restarts is 5; bernstein stop was needed to get the server off the port.

Where it comes from. The orchestrator's quiescence path exits the process on purpose (core/orchestration/orchestrator.py, the "Quiescence confirmed ... self-stopping" branch). The watchdog in core/orchestration/bootstrap.py:955-1020 treats any spawner that is not positively alive as crashed and restarts it, by design, with max_restarts = 5 and restart_reset_after_s = 120.0 (bootstrap.py:1135-1136). Nothing tells the watchdog that this particular exit was the run finishing. Its docstring explains why it must restart a dead spawner; it does not consider a spawner that stopped because there was nothing left to do.

Consequences.

  1. bernstein trace export --last selects "the newest finished run", which is now one of the empty ones, so the export of a finished run has to be done by run id.
  2. Five spurious run_started/run_completed pairs per finished run in the journal history, manifests and retrospectives.
  3. The task server stays up on its port after the run is over.

Also observed, same run: --wait returned before the worker finished

--wait printed its summary at 20:00:55 with Done: 1, Active agents: 1, Elapsed: 0s, and the CLI exited 0. The worker was reaped at 20:01:00 and its branch merged at 20:01:01 (spawner.log lines 82-87). So the summary reported a finished run while the only agent was still running and its work had not been merged. cli/run_preflight.py:741 calls _wait_for_run_completion and then _show_run_summary; whatever _wait_for_run_completion waited on had already gone terminal at that point, although the task (7f8f9bdabf7a) was still claimed by a live agent. Slice 2 below.

What must become true

  1. A quiescence self-stop is not a crash: after it, the watchdog does not restart the spawner and no new run directory is created.
  2. A spawner that dies any other way is still restarted exactly as today. The docstring's reasons hold.
  3. After a finished run, bernstein trace export --last selects the run that did the work.
  4. --wait returns only when no agent is alive and the run's terminal event has been journaled; its summary never says Active agents: N with N > 0.

Where the code is

What Path
Quiescence self-stop src/bernstein/core/orchestration/orchestrator.py:2603-2611 ("Quiescence confirmed ... self-stopping")
Watchdog restart decision src/bernstein/core/orchestration/bootstrap.py:955-1020 (_check_and_restart style helper), :1135-1136 (max_restarts, restart_reset_after_s)
Pidfile liveness classification shared with the CLI classify_pidfile_liveness, referenced from bootstrap.py:985-990
--wait path src/bernstein/cli/run_preflight.py:728-744
--last selection src/bernstein/cli/commands/export_cmd.py

Slices

Slice 1 (take this one alone). The watchdog recognises a deliberate stop. The orchestrator leaves a marker when it self-stops on quiescence (a file next to the pidfile, or a recorded exit reason the liveness classifier reads); the watchdog reads it and stands down instead of restarting. The marker is removed on the next real start. Named test: test_watchdog_does_not_restart_after_quiescence_self_stop (load-bearing, must fail on main), test_watchdog_still_restarts_a_crashed_spawner.

Slice 2. --wait and the run-completion verdict: do not report terminal while a claimed task has a live agent. Named test: test_wait_does_not_return_while_an_agent_is_alive.

Out of scope

  • Changing when the orchestrator decides it is quiescent.
  • The run directory naming, the WAL replay, bernstein stop.
Brief for a coding agent
Repo: sipyourdrink-ltd/bernstein. Do slice 1 only.

The defect: after the orchestrator exits on "Quiescence confirmed ... self-stopping", the watchdog in
core/orchestration/bootstrap.py (:955-1020, max_restarts=5 at :1135) sees a dead spawner pidfile and
restarts it, five times, each start creating an empty .sdd/runs/<id>/ with run_started+run_completed.

Order of work:
1. Write test_watchdog_does_not_restart_after_quiescence_self_stop in the tests next to the existing
   watchdog tests for bootstrap.py: simulate a spawner that recorded a deliberate stop, run one
   watchdog check, assert restart_fn was NOT called. Run it, show it FAILS on main. Paste the failure.
2. Have the orchestrator record the deliberate stop at the self-stopping site (a small marker file in
   .sdd/runtime next to spawner.pid is enough; name it so `bernstein status` can explain it).
3. Make the watchdog read that marker before deciding "dead, restarting"; clear it on the next start.
4. Add test_watchdog_still_restarts_a_crashed_spawner: no marker, dead pid -> restart_fn called.
5. Nothing else: not the quiescence rule, not --wait (slice 2), not WAL replay.

Verify:
  uv run pytest tests/unit/core/orchestration -q -k "watchdog or bootstrap"
  uv run ruff check src tests && uv run ruff format --check <the files you touched>

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

P2Priority 2 - NormalbugSomething isn't workingsize/s

Type

No type

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions