Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
68 changes: 57 additions & 11 deletions docs/backends.md
Original file line number Diff line number Diff line change
Expand Up @@ -116,12 +116,39 @@ ones adapters get wrong:
and Docker Desktop the image's own `USER` runs. On a daemon that enforces bind-mount ownership (native Linux
Docker, rootful Podman) the worker runs the job as its own uid with `--user` and `HOME=/home/pi`, because
only the uid that owns the job's files can use them (a uid-1001 worker, the image's own uid, needs no flag).
The worker decides this from facts at boot and before each job, never by starting a probe container. It
refuses by name what no uid can serve: rootless Docker or Podman, userns-remap, a root worker, Docker
Desktop on Linux (WSL is not affected), an image without the `anyUid` capability for another uid, and a
`--user` whose primary group is 0 or the docker socket's. `pi-dispatch doctor` names the answer for the shell
it runs in, and `pi-dispatch doctor --live` runs its probe as that user and reads it back. An adapter for
another runtime answers the same question in its own terms.
The worker decides this from facts at boot and before each job, never by starting a probe container, and
refuses by name what no uid can serve. **When that refusal fires depends on which cause it is**, so both
sets are named here rather than one:

<!-- worker/test/backends-doc.test.mjs GENERATES the two list lines below from `JOB_USER_FIX` and
`jobUserBootRefusal` and requires each verbatim, so edit them by pasting what that test prints when it
fails. Each
cause name may appear exactly once on this page; to say more about one, say it without the name. -->
- **Stops the boot**: `rootless`, `userns-remap`, `worker-is-root`, `desktop-linux-userns`. Nothing a job
may run as works, so the worker exits rather than picking up work it could only refuse. Three of these
are read off the daemon, though one of them can instead be inferred from the socket this worker was
pointed at; the root-worker case is a fact about the account the worker itself runs as, and is fixed by
changing that account rather than the host. Docker Desktop is refused only on Linux outside WSL, and
WSL2 is not affected. That exit
is conditional, and the condition is real rather than decorative: it happens while `local` is the default
venue (`BOOT_REFUSING_JOB_USER_CAUSES` in `worker/src/job-user.mjs`, read by `jobUserBootRefusal` in
`start.mjs`), and a deployment whose default venue was elsewhere would boot and refuse each local job
instead. `local` is the only venue this build has, so every deployment today gets the exit.
- **Refuses each job**: `runtime-unreadable`, `root-group`, `docker-group`, `any-uid-unsupported`. The
worker boots, and each local job returns a policy refusal naming the cause. The first of those is a
daemon whose answer cannot be read at all, which is not the same as one that has not answered yet.
The last three are on the `--user` path only, so a worker that is already uid 1001, the image's own
uid, meets none of them: nothing is passed, and its primary group is not compared.

A daemon that has not answered YET is in neither set. The decision is `unknown`, it is never cached and never a
boot exit, so a unit carrying `RestartPreventExitStatus=2` is not stranded by a daemon that is still
starting. A job picked up while the answer is still unknown is an infrastructure **retry**, not a refusal
(`CONST-RETRY-INFRA-ONLY`): the processor throws, so the queue tries again once the daemon answers, where a
policy refusal returns and is final.

`pi-dispatch doctor` names the answer for the shell it runs in, and `pi-dispatch doctor --live` runs its
probe as that user and reads it back. An adapter for another runtime answers the same question in its own
terms.

## A declaration is not a claim that the property holds

Expand Down Expand Up @@ -241,17 +268,36 @@ The harness cannot detect that. A green run is not a conformant backend.
## Registering it

Pass your bundle to `startWorker` as an extra backend. It is registered after `local`, and its own `reap`
joins the boot sweep automatically. **There is no published import for `startWorker` today**: it lives in
`worker/src/start.mjs`, which the package's export map does not name, so this works from a checkout of this
repository and not from an installed package. Its first argument is the environment, and the backends ride
the second:
joins the boot sweep automatically. Its first argument is the environment, and the backends ride the second:

```js
import { startWorker } from "./worker/src/start.mjs"; // from a checkout; not an export of @edgehero/pi-dispatch
import { startWorker } from "./worker/src/start.mjs"; // from a checkout of this repository

await startWorker(process.env, { extraBackends: [myBackend] });
```

**Registration is in-tree, and that is a decision rather than a gap** (issue #342). `startWorker` lives in
`worker/src/start.mjs`, which the package's export map does not name, so this works from a checkout and not
from an installed package. Exporting it was considered and refused: the second argument is the worker's
whole dependency-injection bag, test seams included, and publishing an import for it makes every one of
those seams a public API that a release has to keep.

It costs an adapter author nothing they were not already paying, which is the part worth being clear about.
Step 1's `BACKENDS_TABLE` entry is in-tree by necessity, for the reason given at the top of this page: the
declaration is what an operator reads, so a venue that runs jobs without one would be a venue nobody can
reason about. An adapter that already has to land a declaration here loses nothing by registering here too.
The code itself can still live anywhere.

Issue #342 asked for this to be decided alongside the Podman route rather than after it, and it is: issue
#354's `podman` backend is specified in-tree, declaring its own words in the backend table and implementing
the adapter contract on this page, so it needs no export and nothing here blocks it.

Reopening this needs two things, not one, and the second is the harder. An export would let an
npm-installed operator CALL `startWorker`, and it would still leave them unable to register anything:
`BACKENDS` is frozen at module load, `parseBackendList` refuses a `PI_BACKENDS` name it does not know, and
`validateBackend` refuses a trigger naming one. So an out-of-tree venue also needs a way to DECLARE itself,
which is the thing the paragraph above refuses on purpose. Nobody has asked for either.

`startWorker` builds the registry itself:

```js
Expand Down
14 changes: 11 additions & 3 deletions docs/egress.md
Original file line number Diff line number Diff line change
Expand Up @@ -188,9 +188,17 @@ All of it was run. The method costs nothing and is worth repeating on your own h
peer on another job network is unreachable by name and by address. See `docs/podman.md`.
- **A denied host fails in about 20 ms, not on a DNS timeout**, because the client hands the name to the
proxy in a `CONNECT` and never resolves it locally. An external name resolved *directly* from an internal
network fails without ever reaching the proxy, and how fast depends on the resolver in front of the container,
not on this design: about 10 seconds on the host this was first measured on, and under 5 ms in a Linux lab where
nothing answers at all. Either way it is the path a client that bypassed the proxy would take.
network fails without ever reaching the proxy, and how fast depends on the resolver **inside** the
container rather than on anything this design does. Measured on Docker Desktop 27.4.0 (macOS), an
`--internal` network, and docker's embedded resolver at `127.0.0.11`, whose upstream on that host is the
Desktop VM's own resolver and is not reachable from a network with no route off it. So the forward fails
rather than hanging, and what each client does with that failure is the whole of the difference: the job
image, which is Debian and glibc, gives up in about 10 ms (six runs, 6 to 23 ms, all `EAI_AGAIN`), taking
the embedded resolver's answer as final, while a musl image on the same network waits out musl's own 5 s
resolver timeout instead (five runs, 5013 to 5025 ms). The number that describes a job here is therefore
milliseconds, since the job image is glibc. The ten-second figure this replaces did not reproduce on this
host in any client, and what it was measured against is not known, so it is replaced rather than accounted
for. Either way it is the path a client that bypassed the proxy would take.

## Appendix: a host-firewall layer below docker's rules

Expand Down
Loading
Loading