Skip to content

fix(macos): couple forge3 lifecycle to the app via a guardian process - #40

Closed
ssddOnTop wants to merge 2 commits into
mainfrom
fix/couple-app-forge3-lifecycle
Closed

ssddOnTop wants to merge 2 commits into
mainfrom
fix/couple-app-forge3-lifecycle

Conversation

@ssddOnTop

Copy link
Copy Markdown
Collaborator

The bug

Orphaned forge3 processes accumulate. Locally I found four with PPID=1, aged 9 minutes to 2 days, each still LISTENing on a loopback port (9754/9756/9759/9760) while the menu bar app was not running.

The log signature of an orphan is a Started + Stopping pair with nothing after it:

[INFO] Started forge3 process group 1163 ... on private loopback port 9759
[INFO] Stopping forge3 process group 1163
<<< nothing further, ever >>>

versus a healthy stop, which always produces three lines:

[INFO] Stopping forge3 process group 11315
[WARN] forge3 did not terminate within 3.0 seconds; sending SIGKILL
[INFO] forge3 exited with status 137 after 7.8 seconds

Root cause

ForgeProcessHost.stop() sent kill(-pid, SIGTERM) and scheduled the SIGKILL escalation 3 seconds later on its serial queue. Three facts combine:

  1. forge3 never honours SIGTERM. All 652 healthy stops in the local log required the SIGKILL escalation — 652/652. So the entire kill depends on the app staying alive for 3 more seconds.
  2. forge3 runs in its own process group (POSIX_SPAWN_SETPGROUP + posix_spawnattr_setpgroup(&attr, 0)), so it receives no group- or session-scoped signal when the app dies. macOS has no PR_SET_PDEATHSIG equivalent. This was chosen so the app can kill(-pid, …) forward to clean up descendants, but it severs the kernel's backward safety net.
  3. If the app dies inside that 3-second window — force quit, crash, logout, Sparkle relaunch — the pending escalation dies with it.

Nothing swept orphans at startup, and LoopbackEndpoint.allocate() silently bind-probes past occupied ports, so the leak never surfaced. The port number was effectively a leak counter.

Two plausible-looking theories were investigated and ruled out by the logs: the reapExitedProcessLocked early-return hang (zero Termination watchdog expired and zero Could not reap lines), and restart churn invalidating a pending escalation (startLocked guards on pid == 0, and pid is only cleared after waitpid succeeds, so a new spawn cannot pre-empt a live process's escalation).

The fix

ForgeLifecycleGuardian — the same binary re-executed in an internal mode. The app spawns the guardian; the guardian spawns forge3 into the guardian-owned process group; the app holds a framed control socket.

Closing that socket — which the kernel does unconditionally on app death, including SIGKILL — makes the guardian terminate and reap the entire group. Conversely the guardian exits as soon as forge3 exits, so the app sees an ordinary child exit.

Group kills are ordered against PID reuse by observing forge3 with waitid(..., WNOWAIT), which pins the PGID until the single authoritative reap.

This is strictly stronger than the previous in-app escalation, because it no longer depends on the app surviving any window.

The invariant

app dies    -> forge3 dies   (except uncatchable kills of the guardian itself)
forge3 dies -> app survives and restarts it

The restart half is a deliberate divergence. The branch this work is ported from also made a forge3 failure terminate the menu app: it deleted scheduleRestart entirely and gutted restartAfterReadinessFailure into stopAfterReadinessFailure, with AppDelegate calling NSApp.terminate(nil) on unexpected exit.

That policy is not adopted here. ServiceSupervisor.swift is deliberately unchanged — it keeps full ownership of restart behaviour, including the exit-status-75 self-update handshake, which the fail-fast policy would have turned into an app quit. ServiceController likewise no longer escalates a failed phase to termination.

Worth flagging for review: this conflict is semantic, not textual. The two sides edited different methods toward opposing goals, so a plain merge would have produced a clean result with scheduleRestart surviving but no longer called.

Verification

  • swift build clean; swift test 150 tests, 1 skipped, 0 failures.
  • testGuardianControlEOFKillsServiceGroupWithinBound models app crash via control-channel EOF against a trap '' TERM process with a descendant, and asserts both die within 1.5s.
  • testGuardianFailureKillsPersistentLeaderAndDescendantAndReportsUnexpectedExit covers guardian SIGKILL.
  • The existing restart tests still pass unchanged: testUnexpectedExitSchedulesRestartWithNewEndpoint, testUpdateInstalledExitRestartsImmediately, testRepeatedUpdateInstalledExitUsesBackoff.
  • scripts/test-packaging.sh passes.
  • Confirmed assemble-app.sh installs only the app executable into Contents/MacOS/, so the guardian's dev-only ForgeRuntimeLeaseTestHelper preference cannot be reached in a shipped bundle; production always falls through to Bundle.main.executableURL.

Removed

testAppEntrypointLifecycleFailureTerminationProbeExitsPromptly and the --forge-internal-lifecycle-termination-probe entry point existed solely to assert the fail-fast policy, so they were dropped with it. A comment in ProcessIntegrationTests records why.

Not addressed

Pre-existing orphans are not swept at startup — this change prevents new ones but does not reclaim the four already running. A startup sweep would need PID-reuse validation (boot UUID + start time + executable path) to be safe, which is worth doing separately.

ssddOnTop and others added 2 commits August 3, 2026 11:52
Orphaned forge3 processes accumulated: instances with PPID=1 survived the
menu bar app indefinitely, each still LISTENing on a loopback port (four
observed locally, aged 9 minutes to 2 days, on ports 9754/9756/9759/9760).

Root cause. `ForgeProcessHost.stop()` sent `kill(-pid, SIGTERM)` and
scheduled the SIGKILL escalation 3 seconds later via `queue.asyncAfter`.
forge3 never honours SIGTERM -- every one of the 652 healthy stops in the
local log required the escalation -- so the kill depended entirely on the
app surviving those 3 seconds. forge3 is also spawned into its own process
group (POSIX_SPAWN_SETPGROUP), so it receives no group- or session-scoped
signal when the app dies, and macOS has no PR_SET_PDEATHSIG equivalent. If
the app died inside that window (force quit, crash, logout, Sparkle
relaunch), the pending escalation died with it and forge3 leaked. Nothing
swept orphans at startup, and the port allocator silently bind-probed past
them, so the leak was invisible.

The log signature is an orphan having `Started`/`Stopping` lines with no
following `sending SIGKILL` or `exited with status`.

Fix. Introduce `ForgeLifecycleGuardian`: the same binary re-executed in an
internal mode. The app spawns the guardian, which spawns forge3 into the
guardian-owned process group, and the app holds a framed control socket.
Closing that socket -- which the kernel does unconditionally on app death,
including SIGKILL -- makes the guardian terminate and reap the whole group.
The guardian in turn exits as soon as forge3 exits, so the app observes a
normal child exit. Group kills are ordered against PID reuse by observing
forge3 with waitid(WNOWAIT) so the PGID stays pinned until the sole reap.

This makes the coupling one-way, which is the intended invariant:

  app dies      -> forge3 dies (except uncatchable kills of the guardian)
  forge3 dies   -> app survives and restarts it

The restart half is deliberate. The branch this work is ported from also
made a forge3 failure terminate the menu app, deleting `scheduleRestart`
and gutting `restartAfterReadinessFailure`. That policy is not adopted:
`ServiceSupervisor` is unchanged and keeps full ownership of restart
behaviour, including the exit-status-75 self-update handshake, which that
policy would have turned into an app quit. `ServiceController` likewise no
longer escalates a failed phase to termination.

Verified with `testGuardianControlEOFKillsServiceGroupWithinBound`, which
models app crash via control-channel EOF against a SIGTERM-ignoring process
with a descendant, and asserts both die within 1.5s -- alongside the
existing restart tests, which still pass unchanged.
@ssddOnTop ssddOnTop added the test-release Run the full Release workflow as a dry run on this PR (artifact only, no release touched) label Aug 4, 2026
@ssddOnTop

Copy link
Copy Markdown
Collaborator Author

fixed in #45

@ssddOnTop ssddOnTop closed this Aug 6, 2026
@ssddOnTop
ssddOnTop deleted the fix/couple-app-forge3-lifecycle branch August 6, 2026 08:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

test-release Run the full Release workflow as a dry run on this PR (artifact only, no release touched)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant