Repository navigation
fix(renderer): clear the stuck daemon updating state after restart - #38
Merged
Merged
Conversation
added 7 commits
September 30, 2026 12:54
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
After the user clicks the daemon update in the environment window, the top bar shows a "Daemon updating" chip and never clears it. Observed on 2026-09-30: the update was requested at 10:28:36 UTC, and two hours later the chip still read "Daemon updating" while the environment was otherwise working.
Evidence from the logs:
instance.daemon-upgrade {"mode":"now"}at 10:28:36 and nothing after it about the upgrade.daemon-closedat 10:28:36 and reopened at 10:28:37; the control channel closed with reasonapp-closedat 10:29:36, exactly 60 seconds later.daemon.upgrade {"mode":"now"}anddaemon.exit {"code":75}at 10:28:36, thendaemon.startanddaemon.listeningon the new version0.1.0+d48a3a42570danddaemon.readyat 10:28:37, thenclient.attached {"protocol":1,"replay":14}.So the daemon actually upgraded and restarted within a second. The app is what stays stuck. In
src/renderer/instance-store.ts,upgradingis set by thedaemon.upgradingevent and cleared only when a snapshot is applied, and the reattach after the restart replayed 14 events rather than resyncing with a snapshot, so nothing cleared it. Later periodic resync attachments in the container log did not clear it either, so check whether they belong to the app at all, and whether the app's own daemon-version display (state.daemon) also stays stale after the restart.Expected: once the daemon is back on the new version, the chip disappears, the top bar shows the new daemon version, and if the restart never completes the user sees an error instead of an endless "updating".
What Changed
The periodic resync attachments in the container log belong to the runner's GitHub token pump (
src/puck-runner/daemon-link.ts), whose connections do not forward those snapshots to the app's renderer.Risk Assessment
✅ Low: A new-build welcome or snapshot clears the updating chip and refreshes the daemon version, a restart that never finishes becomes an error after two minutes, and a same-build reattach keeps a drain pending only while its original turns are still running.
Testing
I launched the isolated Puck window and drove the top bar through the reported update, a replay onto the new daemon, a two-minute stall, and the drain reconnect cases. The updating chip cleared to 0.1.0+d48a3a42570d, a stall showed Daemon update failed with Retry update, and a drain stayed pending across a short drop and a long same-build replay until its original turns finished.
Evidence: Top-bar text after each step
After replay: chips empty, Daemon 0.1.0+d48a3a42570d (d48a3a4). Stall while attached: Daemon update failed, buttons [Retry update]. Unreachable: buttons [Reconnect]. Drain after a one-second drop and two minutes: Daemon updating, title Updating once running turns finish. Long disconnect: Reconnecting plus Daemon update failed, no buttons; same-build welcome restores Daemon updating; two minutes after the original turns end: Daemon update failed, Retry update.Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 2 issues found → auto-fixed (2) ✅
src/renderer/instance-store.ts:401- A drain that stays disconnected for the two-minute deadline is labeled failed after the app reattaches, even though the same build is back and the original turns are still running.watchUpgrade(line 175) arms the deadline on any non-attached upsert. When it fires while still disconnected, the callback at lines 175-180 clearsupgradingand setsupgradeErrorbut leavesupgradePhaseasdraininganddrainTurnsintact. Coming back is a replay, not a snapshot:instance-sync.ts:153only reads a snapshot when the projection is empty, andinstance-sync.ts:160then callsapplyWelcome. A same-build welcome (lines 403-408) updates the daemon and callswatchUpgrade, and becauseupgradingis already null it returns without undoing the failure. The restore at lines 387-392 runs only insideapplySnapshot, so it never sees this reattach. The chip and banner stay on "Daemon update failed" for the rest of the drain, and Retry (topbar.ts:323) callsdaemon.upgradewhile that process is still inupgradingand is rejected. Concrete sequence: drain with turn T in flight, the attach stays down for two minutes (the turn is still running on that daemon), then the channel returns within the event log so the welcome is a replay of the same build and T is still ininflight. Restore that drain on the same-build welcome with the same condition as lines 387-392 (phase stilldraining, original turns still in flight, status notstopping), so a laterturn.endor a real restart can arm a fresh deadline. Leave the error in place when the daemon stays disconnected.src/renderer/topbar.ts:331- The daemon-update-failed banner always adds a Reconnect button (view.retry || upgradeError), including when the environment is already attached and Retry is the failure action. No intent requirement needs that control: the required outcome is that a restart which never finishes is visible as an error, and the attach view already shows Reconnect when the connection itself can be retried. Recommend removing the|| upgradeErrorbranch so Reconnect stays on the existing attach-view rule.🔧 Fix applied.
1 warning still open:
src/renderer/instance-store.ts:401- A drain that stays disconnected for the two-minute deadline is labeled failed after the app reattaches, even though the same build is back and the original turns are still running.watchUpgrade(line 175) arms the deadline on any non-attached upsert. When it fires while still disconnected, the callback at lines 175-180 clearsupgradingand setsupgradeErrorbut leavesupgradePhaseasdraininganddrainTurnsintact. Coming back is a replay, not a snapshot:instance-sync.ts:153only reads a snapshot when the projection is empty, andinstance-sync.ts:160then callsapplyWelcome. A same-build welcome (lines 403-408) updates the daemon and callswatchUpgrade, and becauseupgradingis already null it returns without undoing the failure. The restore at lines 387-392 runs only insideapplySnapshot, so it never sees this reattach. The chip and banner stay on "Daemon update failed" for the rest of the drain, and Retry (topbar.ts:323) callsdaemon.upgradewhile that process is still inupgradingand is rejected. Concrete sequence: drain with turn T in flight, the attach stays down for two minutes (the turn is still running on that daemon), then the channel returns within the event log so the welcome is a replay of the same build and T is still ininflight. Restore that drain on the same-build welcome with the same condition as lines 387-392 (phase stilldraining, original turns still in flight, status notstopping), so a laterturn.endor a real restart can arm a fresh deadline. Leave the error in place when the daemon stays disconnected.🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
node ~/.no-mistakes/evidence/01M3S4B93NMVHBA3MHMEEF3YJC/drive-daemon-update.mjsagainst the isolated app (PUCK_FIXTURE=full, CDP on the environment window)Clicked Daemon update, chose Now, then delivered a new-build welcome and a 14-event replay including an old upgrading eventAdvanced the window clock past two minutes for an immediate update that never returned, then clicked Retry updateSent a same-build welcome during an immediate update, then advanced two minutesFailed an update while the runner was unreachable and checked the banner actionsDrained with four live turns, dropped the attach for one second, returned on the same build, advanced more than two minutes, then welcomed the new buildEnded those original turns, started a new turn, and advanced two minutesLeft a drain disconnected for two minutes, confirmed the failure stayed, then restored it with a same-build welcome and no snapshot and advanced two minutes after the original turns ended✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.