Skip to content

feat(proxy): port_holder handoff, JSON-verified reboots, LB drift detection (#93 workstream E) - #97

Merged
mhenrixon merged 2 commits into
dashfrom
feat/auto-reboot-completion
Aug 3, 2026
Merged

mhenrixon merged 2 commits into
dashfrom
feat/auto-reboot-completion

Conversation

@mhenrixon

Copy link
Copy Markdown
Collaborator

Summary

Workstream E of the v3.0.0 release gate (#93) — auto-reboot completion. Two commits: the ported port_holder handoff, then its completion.

  • E1 — port_holder handoff ported onto today's dash. Cherry-pick of the PR feat(proxy): zero-downtime generation handoff via port-holder #25 commit (feat/proxy-auto-reboot was stacked pre-R6; the rest of that branch already landed via other PRs). With proxy: run: port_holder: true, a minimal long-lived kamal-proxy-net container owns the published ports; generations join its network namespace with --reuse-port, and a drift-detected reboot becomes run-next → wait-ready → drain-old → wait → remove → rename — no listener ever goes dark. Conflicts resolved against B–D: reuse-port is forced through the existing server_options key (run_config["reuse_port"] || port_holder?), and the holder's publish-suppression composes with the secrets/docker-socket args added since.
  • E2 — verify_services via kamal-proxy list --json. Exact key membership on the JSON services map replaces the substring match over the human table, which could find "app" inside "other-app" and call a missing route verified. re_register_services now returns the registered names.
  • E3 — LB drift detection/auto-reboot. Boot compares the LB container's config-digest label with the current config: drifted + reboot_on_deploy → reboot with pre/post-loadbalancer-reboot hooks; drifted + reboot_on_deploy: false → stale warning, old container keeps serving. The reboot sequence is extracted into Kamal::Cli::Proxy::LoadbalancerReboot, shared verbatim with kamal proxy reboot so the two paths cannot diverge.

E4 (the handoff integration test against the published image) runs in workstream G with bin/test — reboot_on_deploy: true stays the default pending that run, per the settled decision.

Refs #93 (workstream E — F/G remain)

Test plan

  • port_holder suite from PR feat(proxy): zero-downtime generation handoff via port-holder #25 green against B–D-era code (run command, holder boot, handoff sequence, migration path, digest adoption)
  • list --json command shape; reboot fails on a missing service; completes only when the JSON listing contains it (the generic stub fails JSON.parse, proving the path)
  • Drifted LB reboots on boot; reboot_on_deploy: false warns and keeps the old container serving
  • LB still publishes 80/443 in port_holder mode
  • bundle exec rubocop --parallel clean; unit suite green (builder failures are the known Apple-Silicon artifacts)
  • CI; E4 handoff integration test in workstream G

Deviations & judgment calls

  • Port, not merge: feat/proxy-auto-reboot was not merged wholesale — most of it reached dash long ago through other PRs, so E1 is a cherry-pick of the one missing commit (e79e6c8a), conflict-resolved against the B–D-era files.
  • Discovery — LB × port_holder interaction: the shared run surface (from feat(proxy): explicit loadbalancer layering contract + LB run plumbing (#93 workstream B) #94) suppresses publish_args in holder mode, which would have left a dedicated LB publishing nothing (it has no holder on its host). Loadbalancer#run_args now re-adds publishing in holder mode; the LB's reboot path stays stop→run. The LB digest intentionally still matches the proxy digest (it hashes the shared surface) — it is only ever compared against itself.
  • Judgment call: port_holder remains opt-in (proxy/run/port_holder: true), as PR feat(proxy): zero-downtime generation handoff via port-holder #25 shipped it. F3's release-notes item reads as if the post-upgrade fleet reboot is holder-mode by default — if you want port_holder defaulting true for 3.0, that's a one-line flip plus a digest note, but I didn't make that call unilaterally.
  • In-path fix: D4 missed cutting lease_ttl/lease_wait from docs/proxy.yml — since the docs file is the validation schema, the keys still validated while emitting nothing. Cut here (found while resolving the docs conflict in that exact region).
  • Refactor: verify_services takes the registered list as a parameter instead of reading an ivar — testable seam, no behavior change.
  • JSON parse failures raise (no rescue): a proxy whose list --json returns non-JSON after a reboot is a real failure, not something to paper over.

…run port_holder) (#25)

With proxy: run: port_holder: true, a minimal long-lived kamal-proxy-net
container (same image, running kamal-proxy hold) owns the published
ports and the kamal network attachment. Proxy generations join its
network namespace (--network container:kamal-proxy-net) and run with
--reuse-port, so two generations can hold :80/:443 simultaneously.

A drift-detected reboot then becomes a handoff instead of a gap:

  run kamal-proxy-next (restores state, binds alongside the old
  generation) -> wait until its control socket answers -> docker exec
  kamal-proxy drain (old finishes in-flight, flushes state, exits) ->
  docker wait -> remove old -> rename kamal-proxy-next to kamal-proxy

The kernel balances new connections across both generations during the
overlap; no listener ever goes dark. Deploy lock is held throughout, so
no service mutations race the state file.

Adoption is automatic and self-healing: flipping port_holder changes the
config digest, so the next deploy reboots — the holder cannot bind ports
still held by a pre-holder proxy, so that one migration takes a final
brief gap. Fresh hosts boot the holder before the generation
(start_holder_or_run). Requires a kamal-proxy image with the hold,
drain, and --reuse-port commands.

Legacy boot_config deployments are unaffected (port_holder requires a
proxy/run block).
## Summary

Completes workstream E of #93 on top of the ported port_holder handoff.

E2: Reboot#verify_services parses `kamal-proxy list --json` and checks exact
service-name membership - the plain-text substring match could find one
service's name inside another's ("app" in "other-app") and call a missing
route verified. re_register_services now returns the registered names
instead of stashing them in an ivar.

E3: the auto-activated loadbalancer gets the same drift detection the proxy
hosts have. Boot compares the running container's config-digest label with
the current configuration: drifted + reboot_on_deploy reboots it (with
pre/post-loadbalancer-reboot hooks); drifted + reboot_on_deploy: false warns
and leaves the old container serving. The reboot sequence is extracted into
Kamal::Cli::Proxy::LoadbalancerReboot, shared verbatim by `kamal proxy
reboot` so the two paths cannot diverge.

Also: the loadbalancer keeps publishing its own ports when the proxy fleet
runs in port_holder mode - it has no holder on its host (its reboot path is
stop -> run), and the holder-mode publish suppression from the shared run
surface would have left it listening on nothing. And the lease_ttl /
lease_wait docs entries missed by the D4 cut are gone - the docs file is the
validation schema, so they validated while emitting nothing.

## Test Coverage

- list --json command shape; missing-service reboot failure; JSON-verified
  success (the generic stub would fail JSON.parse, proving the path)
- drifted LB reboots on boot; reboot_on_deploy: false warns and keeps the
  old container serving
- LB publishes 80/443 under port_holder mode

## Verification

- [x] bundle exec rubocop --parallel passes
- [x] unit suite passes (builder failures are the known host-arch artifacts)

Refs #93
@mhenrixon mhenrixon self-assigned this Aug 3, 2026
@mhenrixon
mhenrixon merged commit 01a2cfd into dash Aug 3, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant