Skip to content

Epic: dash / dash-proxy improvement roadmap (R1–R5) #13

Description

@mhenrixon

Tracking epic for the dash / dash-proxy improvement roadmap — bridging selected gaps vs nginx/traefik/caddy/envoy for kamal's audience, without competing head-on. Full rationale and evidence links live in ROADMAP.md in each repo.

Strategy: own what basecamp rejected (headers, rate limiting, compression, PROXY protocol), port what they left stuck (on-demand TLS, mTLS, basic auth, scale-to-zero), ship the table-stakes resilience knobs every other proxy has.

Structure: child issues live in whichever repo the code changes — proxy work in mhenrixon/kamal-proxy, gem work here. Features needing both get one issue per repo, cross-linked. Grouped by R1–R5 milestones.

Release ordering (hard constraint): the proxy image ships BEFORE the gem in any release where MINIMUM_VERSION moves — the gem's minimum must name a published ghcr tag. Versioning: gem 2.12.0.N, proxy v0.9.2.N. See .claude/rules/upstream-sync.md.

R1 — Foundations & fixes

All S-sized: fix real bugs and expose already-built capability (rollout!). Max value per line. Ships as gem 2.12.0.1 + proxy v0.9.2.2.

R2 — Timeouts & resilience

The #1 pain (per-route timeouts) plus the resilience story no kamal-proxy user gets today without fronting nginx/traefik.

R3 — Security & access

Independent S-sized middlewares — good filler between larger efforts. Several are explicit basecamp rejections (safe fork moat).

R4 — TLS & custom domains

The single biggest differentiator (on-demand TLS) — a waiting user base on the stuck upstream PR. Schedule a focused block.

R5 — Traffic shaping & headers

Grab bag ordered by demand: headers, weighted canary, redirects, compression, scale-to-zero, observability.

R6 — Deploy safety for non-proxied roles

Gem-only: no proxy code, so MINIMUM_VERSION does not move and this schedules independently of the proxy release train.

Closes the readiness gap that lets dash stop an old non-proxied container ~7s after the new one merely starts. This is not a Kamal bug — 0efb5ccf ("Remove the healthcheck step") traded it for deploy speed and said so in the commit body, then mitigated it with a barrier that only ever gates on the primary role's healthcheck. A non-primary role's own readiness is still never checked. Realized as a live card decline on a NATS listener role: the old container was SIGTERMed before the new one's subscription existed.

Order matters — #36 first, it unblocks most of the rest.

Suggested landing order: #45 + #43 (visibility, zero behaviour change) → #36 + #39 (the capability) → #37 warn-only + #44 → #38 and #37-as-error behind minimum_version → #40, #41, #42 as needed. #46 and #47 are independent and can go any time.

R7 — kamal-proxy flag coverage

The proxy shipped 23 feat commits since v0.9.2.2 and the gem's config surface never followed. Measured against internal/cmd/:

Surface Proxy accepts Gem exposes Unreachable from deploy.yml
kamal-proxy deploy 80 flags 34 45
kamal-proxy run 31 flags 3 25

The run side is the worse of the two, and for a non-obvious reason: proxy.run.options looks like a passthrough for proxy flags but is consumed by docker_options_args — it passes options to docker run, not to kamal-proxy run. run_command_options emits exactly debug, metrics-port and recheck-targets-on-restore. Everything else is reachable only through a flag's env-var fallback, if it has one.

Net effect: most of R3 (rate limiting, IP allow lists, mTLS), most of R5 (cache, compression, headers, redirects, canary, scale-to-zero) and all of the ACME/DNS-01 configuration are built, tested and shipping in the image — and cannot be turned on from deploy.yml.

Landing order: #82 first — it lands the current gap as issue-linked waivers, which makes every issue after it self-verifying (implement one, delete its waiver, the test proves the flag is really wired). Then #73 (unblocks the DNS epic's value to dash users) and #81 (the run surface, which #73 and #75 both build on). The rest are independent.

Not in this epic: ACME DNS-01 provider coverage in the proxy — zoolutions/dash-proxy#77. That epic adds providers to the Go binary; #73 here makes them configurable. Distinct concerns, distinct repos, no ordering dependency between them.


Release plan

Two tags, ordered — the proxy image must be pullable at the tag MINIMUM_VERSION names before the gem ships.

  1. kamal-proxy → v1.0.0.0 — 57 commits since v0.9.2.2 (38 non-merge: 23 feat, 3 fix), spanning R2–R5.
  2. dash gem → v3.0.0 — MINIMUM_VERSION moves to v1.0.0.0, plus whatever R7 lands.

Known constraint on v1.0.0.0: Gem::Version.new("1.0.0.0") == Gem::Version.new("1.0.0") — trailing zeros are stripped in comparison, so this tag sorts equal to a future upstream v1.0.0 rather than above it. Consequences, both bounded:

  • Kamal::Utils.older_version? would accept a running upstream v1.0.0 proxy as satisfying MINIMUM_VERSION: v1.0.0.0. In practice a dash deploy pulls the ghcr fork image, so this is only reachable if someone points proxy.run.repository at basecamp's.
  • The next fork counter (v1.0.0.1) sorts above both, so the invariant self-corrects on the first patch release.

The git tags themselves never collide. Recorded here so the next person reading .claude/rules/git-workflow.md — which specifies counters starting at .1 for exactly this reason — knows the deviation was deliberate.

dash v3.0.0 also departs from the dash-v<upstream>.<n> gem grammar, declaring a fork major rather than tracking upstream's base. Follow-on: lib/kamal/version.rb will now conflict on every upstream sync, where the playbook currently says "take upstream's". Update .claude/rules/upstream-sync.md as part of the release so the conflict rule matches reality.


Generated from ROADMAP.md. Each issue carries its evidence (basecamp issue/PR numbers) and code anchors.

Activity

  1. added
    epicUmbrella item tracking several smaller issues
    on Jul 3, 2026
  2. 9 remaining items

  3. mhenrixon commented on Jul 11, 2026

    @mhenrixon
    CollaboratorAuthor

    Cutover prioritization is now tracked on the Traefik → dash cutover board (Phase field: P0 Certs → P1 Request chain → P2 Redirects → P3 Cutover ops → Backlog).

    Findings from the edge-config gap analysis (67 features audited against the dash branches):

  4. mhenrixon commented on Jul 29, 2026

    @mhenrixon
    CollaboratorAuthor

    Correction to the R7 table above — the flag counts were understated.

    They came from a regex over internal/cmd/*.go that silently missed every flag registered via Int64Var/Uint16Var, because the method name contains digits. Regenerating from Cobra's own --help (see #83, bin/sync-proxy-flags) against v1.0.0.0:

    Surface Proxy accepts Gem emits Unreachable
    kamal-proxy deploy 80 34 45
    kamal-proxy run 31 3 25

    Missed flags were cache-max-variants (deploy) and cache-lease-ttl / cache-lease-wait (run), all from the cache work in the unreleased set. They are covered by #75. No flag moves between R7 issues — only the totals change.

    Table in the epic body updated.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    epicUmbrella item tracking several smaller issues

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions