Problem
A run that finishes cleanly is restarted five times by the watchdog, and each restart opens a new run directory with nothing in it. Seen on the released build, one goal, one worker:
uvx --from "bernstein[trace]==3.20.0" bernstein run --port 8077 --worker backend --model claude-haiku-4-5-20251001 --auto-approve --quiet --wait
.sdd/runtime/spawner.log and .sdd/runtime/watchdog.log, same run, times local:
20:01:03 Agent backend-99365805 finished task 7f8f9bdabf7a, process reaped
20:01:05 Quiescence confirmed after 2.0s settle window (tick #7, open=0 agents=0, no active holds) - self-stopping
20:01:09 Orchestrator stopped (replay: .sdd/runs/20260920-165918/journal.jsonl, ...)
20:01:12 WARNING Orchestrator (PID 55681) is dead, restarting... <- watchdog.log
20:01:14 Orchestrator started (poll=3s, max_agents=1, server=http://127.0.0.1:8077)
20:01:14 WAL recovery: found 2 uncommitted entries from previous run(s)
20:01:16 Quiescence confirmed after 2.0s settle window (tick #1, open=0 agents=0, no active holds) - self-stopping
20:01:19 Orchestrator stopped (replay: .sdd/runs/20260920-170114/journal.jsonl, ...)
20:01:22 WARNING Orchestrator (PID 59343) is dead, restarting...
... same at 20:01:32, 20:01:42, 20:01:52 ...
Result on disk: .sdd/runs/ holds the real run (20260920-165918, 18 events, run_completed) plus five directories 20260920-1701{14,24,34,44,54} whose journals are run_started, loaded_extension_set, tick_start, task_completed, run_quiescence, run_completed and nothing else. The loop only ends because max_restarts is 5; bernstein stop was needed to get the server off the port.
Where it comes from. The orchestrator's quiescence path exits the process on purpose (core/orchestration/orchestrator.py, the "Quiescence confirmed ... self-stopping" branch). The watchdog in core/orchestration/bootstrap.py:955-1020 treats any spawner that is not positively alive as crashed and restarts it, by design, with max_restarts = 5 and restart_reset_after_s = 120.0 (bootstrap.py:1135-1136). Nothing tells the watchdog that this particular exit was the run finishing. Its docstring explains why it must restart a dead spawner; it does not consider a spawner that stopped because there was nothing left to do.
Consequences.
bernstein trace export --last selects "the newest finished run", which is now one of the empty ones, so the export of a finished run has to be done by run id.
- Five spurious
run_started/run_completed pairs per finished run in the journal history, manifests and retrospectives.
- The task server stays up on its port after the run is over.
Also observed, same run: --wait returned before the worker finished
--wait printed its summary at 20:00:55 with Done: 1, Active agents: 1, Elapsed: 0s, and the CLI exited 0. The worker was reaped at 20:01:00 and its branch merged at 20:01:01 (spawner.log lines 82-87). So the summary reported a finished run while the only agent was still running and its work had not been merged. cli/run_preflight.py:741 calls _wait_for_run_completion and then _show_run_summary; whatever _wait_for_run_completion waited on had already gone terminal at that point, although the task (7f8f9bdabf7a) was still claimed by a live agent. Slice 2 below.
What must become true
- A quiescence self-stop is not a crash: after it, the watchdog does not restart the spawner and no new run directory is created.
- A spawner that dies any other way is still restarted exactly as today. The docstring's reasons hold.
- After a finished run,
bernstein trace export --last selects the run that did the work.
--wait returns only when no agent is alive and the run's terminal event has been journaled; its summary never says Active agents: N with N > 0.
Where the code is
| What |
Path |
| Quiescence self-stop |
src/bernstein/core/orchestration/orchestrator.py:2603-2611 ("Quiescence confirmed ... self-stopping") |
| Watchdog restart decision |
src/bernstein/core/orchestration/bootstrap.py:955-1020 (_check_and_restart style helper), :1135-1136 (max_restarts, restart_reset_after_s) |
| Pidfile liveness classification shared with the CLI |
classify_pidfile_liveness, referenced from bootstrap.py:985-990 |
--wait path |
src/bernstein/cli/run_preflight.py:728-744 |
--last selection |
src/bernstein/cli/commands/export_cmd.py |
Slices
Slice 1 (take this one alone). The watchdog recognises a deliberate stop. The orchestrator leaves a marker when it self-stops on quiescence (a file next to the pidfile, or a recorded exit reason the liveness classifier reads); the watchdog reads it and stands down instead of restarting. The marker is removed on the next real start. Named test: test_watchdog_does_not_restart_after_quiescence_self_stop (load-bearing, must fail on main), test_watchdog_still_restarts_a_crashed_spawner.
Slice 2. --wait and the run-completion verdict: do not report terminal while a claimed task has a live agent. Named test: test_wait_does_not_return_while_an_agent_is_alive.
Out of scope
- Changing when the orchestrator decides it is quiescent.
- The run directory naming, the WAL replay,
bernstein stop.
Brief for a coding agent
Repo: sipyourdrink-ltd/bernstein. Do slice 1 only.
The defect: after the orchestrator exits on "Quiescence confirmed ... self-stopping", the watchdog in
core/orchestration/bootstrap.py (:955-1020, max_restarts=5 at :1135) sees a dead spawner pidfile and
restarts it, five times, each start creating an empty .sdd/runs/<id>/ with run_started+run_completed.
Order of work:
1. Write test_watchdog_does_not_restart_after_quiescence_self_stop in the tests next to the existing
watchdog tests for bootstrap.py: simulate a spawner that recorded a deliberate stop, run one
watchdog check, assert restart_fn was NOT called. Run it, show it FAILS on main. Paste the failure.
2. Have the orchestrator record the deliberate stop at the self-stopping site (a small marker file in
.sdd/runtime next to spawner.pid is enough; name it so `bernstein status` can explain it).
3. Make the watchdog read that marker before deciding "dead, restarting"; clear it on the next start.
4. Add test_watchdog_still_restarts_a_crashed_spawner: no marker, dead pid -> restart_fn called.
5. Nothing else: not the quiescence rule, not --wait (slice 2), not WAL replay.
Verify:
uv run pytest tests/unit/core/orchestration -q -k "watchdog or bootstrap"
uv run ruff check src tests && uv run ruff format --check <the files you touched>
Problem
A run that finishes cleanly is restarted five times by the watchdog, and each restart opens a new run directory with nothing in it. Seen on the released build, one goal, one worker:
.sdd/runtime/spawner.logand.sdd/runtime/watchdog.log, same run, times local:Result on disk:
.sdd/runs/holds the real run (20260920-165918, 18 events,run_completed) plus five directories20260920-1701{14,24,34,44,54}whose journals arerun_started,loaded_extension_set,tick_start,task_completed,run_quiescence,run_completedand nothing else. The loop only ends becausemax_restartsis 5;bernstein stopwas needed to get the server off the port.Where it comes from. The orchestrator's quiescence path exits the process on purpose (
core/orchestration/orchestrator.py, the "Quiescence confirmed ... self-stopping" branch). The watchdog incore/orchestration/bootstrap.py:955-1020treats any spawner that is not positively alive as crashed and restarts it, by design, withmax_restarts = 5andrestart_reset_after_s = 120.0(bootstrap.py:1135-1136). Nothing tells the watchdog that this particular exit was the run finishing. Its docstring explains why it must restart a dead spawner; it does not consider a spawner that stopped because there was nothing left to do.Consequences.
bernstein trace export --lastselects "the newest finished run", which is now one of the empty ones, so the export of a finished run has to be done by run id.run_started/run_completedpairs per finished run in the journal history, manifests and retrospectives.Also observed, same run:
--waitreturned before the worker finished--waitprinted its summary at 20:00:55 withDone: 1,Active agents: 1,Elapsed: 0s, and the CLI exited 0. The worker was reaped at 20:01:00 and its branch merged at 20:01:01 (spawner.log lines 82-87). So the summary reported a finished run while the only agent was still running and its work had not been merged.cli/run_preflight.py:741calls_wait_for_run_completionand then_show_run_summary; whatever_wait_for_run_completionwaited on had already gone terminal at that point, although the task (7f8f9bdabf7a) was still claimed by a live agent. Slice 2 below.What must become true
bernstein trace export --lastselects the run that did the work.--waitreturns only when no agent is alive and the run's terminal event has been journaled; its summary never saysActive agents: Nwith N > 0.Where the code is
src/bernstein/core/orchestration/orchestrator.py:2603-2611("Quiescence confirmed ... self-stopping")src/bernstein/core/orchestration/bootstrap.py:955-1020(_check_and_restartstyle helper),:1135-1136(max_restarts,restart_reset_after_s)classify_pidfile_liveness, referenced frombootstrap.py:985-990--waitpathsrc/bernstein/cli/run_preflight.py:728-744--lastselectionsrc/bernstein/cli/commands/export_cmd.pySlices
Slice 1 (take this one alone). The watchdog recognises a deliberate stop. The orchestrator leaves a marker when it self-stops on quiescence (a file next to the pidfile, or a recorded exit reason the liveness classifier reads); the watchdog reads it and stands down instead of restarting. The marker is removed on the next real start. Named test:
test_watchdog_does_not_restart_after_quiescence_self_stop(load-bearing, must fail onmain),test_watchdog_still_restarts_a_crashed_spawner.Slice 2.
--waitand the run-completion verdict: do not report terminal while a claimed task has a live agent. Named test:test_wait_does_not_return_while_an_agent_is_alive.Out of scope
bernstein stop.Brief for a coding agent