Repository navigation
feat(proxy): port_holder handoff, JSON-verified reboots, LB drift detection (#93 workstream E) - #97
Merged
Merged
Conversation
…run port_holder) (#25) With proxy: run: port_holder: true, a minimal long-lived kamal-proxy-net container (same image, running kamal-proxy hold) owns the published ports and the kamal network attachment. Proxy generations join its network namespace (--network container:kamal-proxy-net) and run with --reuse-port, so two generations can hold :80/:443 simultaneously. A drift-detected reboot then becomes a handoff instead of a gap: run kamal-proxy-next (restores state, binds alongside the old generation) -> wait until its control socket answers -> docker exec kamal-proxy drain (old finishes in-flight, flushes state, exits) -> docker wait -> remove old -> rename kamal-proxy-next to kamal-proxy The kernel balances new connections across both generations during the overlap; no listener ever goes dark. Deploy lock is held throughout, so no service mutations race the state file. Adoption is automatic and self-healing: flipping port_holder changes the config digest, so the next deploy reboots — the holder cannot bind ports still held by a pre-holder proxy, so that one migration takes a final brief gap. Fresh hosts boot the holder before the generation (start_holder_or_run). Requires a kamal-proxy image with the hold, drain, and --reuse-port commands. Legacy boot_config deployments are unaffected (port_holder requires a proxy/run block).
## Summary Completes workstream E of #93 on top of the ported port_holder handoff. E2: Reboot#verify_services parses `kamal-proxy list --json` and checks exact service-name membership - the plain-text substring match could find one service's name inside another's ("app" in "other-app") and call a missing route verified. re_register_services now returns the registered names instead of stashing them in an ivar. E3: the auto-activated loadbalancer gets the same drift detection the proxy hosts have. Boot compares the running container's config-digest label with the current configuration: drifted + reboot_on_deploy reboots it (with pre/post-loadbalancer-reboot hooks); drifted + reboot_on_deploy: false warns and leaves the old container serving. The reboot sequence is extracted into Kamal::Cli::Proxy::LoadbalancerReboot, shared verbatim by `kamal proxy reboot` so the two paths cannot diverge. Also: the loadbalancer keeps publishing its own ports when the proxy fleet runs in port_holder mode - it has no holder on its host (its reboot path is stop -> run), and the holder-mode publish suppression from the shared run surface would have left it listening on nothing. And the lease_ttl / lease_wait docs entries missed by the D4 cut are gone - the docs file is the validation schema, so they validated while emitting nothing. ## Test Coverage - list --json command shape; missing-service reboot failure; JSON-verified success (the generic stub would fail JSON.parse, proving the path) - drifted LB reboots on boot; reboot_on_deploy: false warns and keeps the old container serving - LB publishes 80/443 under port_holder mode ## Verification - [x] bundle exec rubocop --parallel passes - [x] unit suite passes (builder failures are the known host-arch artifacts) Refs #93
This was referenced Aug 3, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Workstream E of the v3.0.0 release gate (#93) — auto-reboot completion. Two commits: the ported port_holder handoff, then its completion.
dash. Cherry-pick of the PR feat(proxy): zero-downtime generation handoff via port-holder #25 commit (feat/proxy-auto-rebootwas stacked pre-R6; the rest of that branch already landed via other PRs). Withproxy: run: port_holder: true, a minimal long-livedkamal-proxy-netcontainer owns the published ports; generations join its network namespace with--reuse-port, and a drift-detected reboot becomes run-next → wait-ready → drain-old → wait → remove → rename — no listener ever goes dark. Conflicts resolved against B–D:reuse-portis forced through the existingserver_optionskey (run_config["reuse_port"] || port_holder?), and the holder's publish-suppression composes with the secrets/docker-socket args added since.verify_servicesviakamal-proxy list --json. Exact key membership on the JSONservicesmap replaces the substring match over the human table, which could find"app"inside"other-app"and call a missing route verified.re_register_servicesnow returns the registered names.reboot_on_deploy→ reboot withpre/post-loadbalancer-reboothooks; drifted +reboot_on_deploy: false→ stale warning, old container keeps serving. The reboot sequence is extracted intoKamal::Cli::Proxy::LoadbalancerReboot, shared verbatim withkamal proxy rebootso the two paths cannot diverge.E4 (the handoff integration test against the published image) runs in workstream G with
bin/test—reboot_on_deploy: truestays the default pending that run, per the settled decision.Refs #93 (workstream E — F/G remain)
Test plan
list --jsoncommand shape; reboot fails on a missing service; completes only when the JSON listing contains it (the generic stub failsJSON.parse, proving the path)reboot_on_deploy: falsewarns and keeps the old container servingbundle exec rubocop --parallelclean; unit suite green (builder failures are the known Apple-Silicon artifacts)Deviations & judgment calls
feat/proxy-auto-rebootwas not merged wholesale — most of it reacheddashlong ago through other PRs, so E1 is a cherry-pick of the one missing commit (e79e6c8a), conflict-resolved against the B–D-era files.publish_argsin holder mode, which would have left a dedicated LB publishing nothing (it has no holder on its host).Loadbalancer#run_argsnow re-adds publishing in holder mode; the LB's reboot path stays stop→run. The LB digest intentionally still matches the proxy digest (it hashes the shared surface) — it is only ever compared against itself.proxy/run/port_holder: true), as PR feat(proxy): zero-downtime generation handoff via port-holder #25 shipped it. F3's release-notes item reads as if the post-upgrade fleet reboot is holder-mode by default — if you wantport_holderdefaultingtruefor 3.0, that's a one-line flip plus a digest note, but I didn't make that call unilaterally.lease_ttl/lease_waitfromdocs/proxy.yml— since the docs file is the validation schema, the keys still validated while emitting nothing. Cut here (found while resolving the docs conflict in that exact region).verify_servicestakes the registered list as a parameter instead of reading an ivar — testable seam, no behavior change.list --jsonreturns non-JSON after a reboot is a real failure, not something to paper over.