Skip to content

fix(podman): read bounds from the user manager, and three venue leftovers (#453) - #456

Merged
edgehero merged 1 commit into
mainfrom
fix/453-podman-venue-leftovers
Sep 28, 2026
Merged

edgehero merged 1 commit into
mainfrom
fix/453-podman-venue-leftovers

Conversation

@edgehero

Copy link
Copy Markdown
Owner

Closes #453.

Bounds: the observation read the wrong fact

podmanBoundsDelegated read podman info's host.cgroupControllers. That list describes the caller's own cgroup, not what is delegated to the account. Measured on Fedora 44 with rootless Podman 5.8.1 and the venue's own argv:

context cgroup manager bounds in the job host.cgroupControllers
linger off, sudo -iu (no user manager) falls back to cgroupfs not applied (max) all five
plain ssh login systemd applied
linger on systemd applied
explicit cgroup_manager="cgroupfs", linger on cgroupfs applied
explicit cgroupfs, linger off cgroupfs not applied
user unit without Delegate= systemd applied (cpu too) memory pids

The bounds were applied exactly when user@<uid>.service ran with pids, memory and cpu delegated to it, whichever cgroup manager Podman used. The observation now reads /sys/fs/cgroup/user.slice/user-<uid>.slice/user@<uid>.service/cgroup.controllers. A missing directory is a determinate miss that names the missing user manager.

Behaviour change: a worker whose PI_BACKEND_FLOOR asks for isolation refuses to boot when no user manager runs. Before this, it ran jobs whose bounds were silently unapplied. Without that floor, jobs are unchanged; doctor warns. The issue's "a plain ssh" framing was wrong: an ssh login starts the user manager, and the failing case is sudo -iu or su with linger off.

Doctor's static line and --live fix, OBSERVATIONS/OBSERVATION_FIX, and the forge comment follow the new rule. Doctor names the cause once.

up, init, doctor

  • up's docker egress step counted an exited proxy as present (exit 0 from docker inspect). It now reads {{.State.Running}} and offers docker start for the shipped proxy, as compose does for an unchanged stopped container. A custom PI_EGRESS_PROXY is reported, never started.
  • init prints the podman ladder when podman is the only venue (PI_BACKENDS from the shell or .env). The docker text is byte-identical.
  • The podman image fix lines no longer suggest docker save ... | podman load.

Specs

DES-PODMAN-NATIVE-ROOTLESS-BACKEND, DES-PODMAN-STACK-AS-QUADLET-UNITS, REQ-DEPLOYMENT-BOOTSTRAP, INT-LIVE-PROBE-CONTRACT; docs/podman.md step 3, step 5 and the property table.

@edgehero
edgehero force-pushed the fix/453-podman-venue-leftovers branch 2 times, most recently from a4a742e to 702d60e Compare September 28, 2026 06:01
…, and three venue leftovers (#453)

Bounds. podmanBoundsDelegated read podman info's host.cgroupControllers,
which is the caller's own cgroup, not what is delegated to the
account: it listed all five controllers with nothing applied, and
memory and pids where cpu was applied. Measured on Fedora 44 / Podman
5.8.1 and Ubuntu 24.04 / 4.9.3 with the venue's own argv, the bounds
apply when user@<uid>.service runs with pids, memory and cpu delegated
and Podman reaches it; with no user manager, or with Podman unable to
reach it over D-Bus (it then falls back to cgroupfs under the caller's
root-owned cgroup), they do not. Credit now needs that cgroup's
controllers and either the systemd cgroup manager or the worker itself
inside user@<uid>.service. Doctor names each cause with its own fix,
reads linger and warns when the manager lives only for a session; the
conformance job carries an isolation floor.

up, doctor and the preflight. A proxy counts as present only when
running, on the pinned digest, with its measured entrypoint and command
(Podman 4.9's string form normalised) and with this folder's two
mounts, each judged on its own (unknown only under a VM path prefix or
on macOS/Windows; on Linux a missing source is stale). A current
stopped proxy is started, a paused one unpaused, a stale one replaced
only with consent (never under --yes while job networks are attached),
quietly reusing the egress network. The worker's pre-spend preflight
admits only a running proxy; paused, exited, dead and created refuse;
any other state is retried once, then fails, naming the proxy and its
state.

doctor reads the deployment's venue keys, VALKEY_URL, PI_PROVIDER and
the provider key from .env through one allowlist and the service's own
loader, says the source, never prints a secret (URLs show only
scheme://host:port/db) and prints fleet names without control bytes;
the podman image fix lines no longer suggest docker save | podman load.

init. The next steps follow up's venue decision: the podman ladder
(linger, pull, up, doctor, service install, doctor --live) on a
podman-only host, none when the venue is unknown, today's docker text
otherwise.

Specs: DES-PODMAN-NATIVE-ROOTLESS-BACKEND, DES-PODMAN-STACK-AS-QUADLET-UNITS,
REQ-DEPLOYMENT-BOOTSTRAP, INT-LIVE-PROBE-CONTRACT, INT-EGRESS-POLICY-CONTRACT.

Signed-off-by: Rob Boerman <robboerman@live.nl>
@edgehero
edgehero force-pushed the fix/453-podman-venue-leftovers branch from a08f482 to 5a637a4 Compare September 28, 2026 06:42
@edgehero
edgehero merged commit d081723 into main Sep 28, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

podman venue leftovers: up's docker proxy check, init's next steps, doctor's cgroup manager line

1 participant