Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 5 additions & 3 deletions .env.example
Original file line number Diff line number Diff line change
Expand Up @@ -114,19 +114,21 @@ VALKEY_URL=redis://127.0.0.1:6379
# On rootful Podman, run these and every docker command through a docker context pointed at podman.sock, with the real docker CLI (docs/podman.md)
PI_JOB_IMAGE=pi-job:latest
# How much memory and CPU each job container gets, unless its project's row in the scoped-limits file sets its own.
# See docs/scoped-limits.md.
# See docs/scoped-limits.md, and docs/sizing.md for how to choose sizes.
# Memory is a whole number of megabytes or gigabytes (512m, 1536m, 4g), at least 512m.
# CPUs is a number with at most two decimals (0.5, 2, 1.25), at least 0.25.
# A bad value stops the worker at boot.
# A job gets no swap beyond its memory where the runtime enforces swap limits (doctor warns where it does not).
# CPUs are a weight, not a cap: under contention a job with more CPUs gets more CPU than one with fewer.
# On an idle host one job may use every core but one (when the host has 4 or more).
# On an idle host one job may use every core up to the CPU budget below, whenever one is in force (with the
# default auto budget, every core but one on a host with 4 or more; with the CPU budget off, the same).
# Default 4g and 2.
# PI_JOB_MEMORY=4g
# PI_JOB_CPUS=2
# How much memory and CPU this host's jobs may hold together: the host budget. A job starts only when its size fits
# beside what already runs here, and a waiting job keeps its place, so a big job is not starved by a stream of small ones.
# PI_CONCURRENCY still caps the number of jobs; whichever is reached first applies. See docs/multi-host.md.
# PI_CONCURRENCY still caps the number of jobs; whichever is reached first applies. See docs/multi-host.md,
# and docs/sizing.md for a worked example.
# Each budget is auto, a value, or off (no limit on that resource).
# auto reads the container runtime: its memory and CPU count (and, on rootless Podman, the user service's own limits),
# minus the reserve below, and never less than one job of the default size above.
Expand Down
7 changes: 6 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,10 @@ What you get:
dollar total per day, week or month with a floor per project, and a portfolio manager flow can move money
between projects inside it with no keypress. Every move is logged and can be reverted
([`docs/projects.md`](docs/projects.md), [`docs/allocation.md`](docs/allocation.md)).
- **A size per project, inside a budget per machine.** Each project's jobs get the memory and CPUs its row sets
(4 GB and 2 CPUs by default). A job starts only when its size fits beside what already runs on that machine, and
room is kept for a waiting big job, so a stream of small ones cannot starve it. `pi-dispatch doctor` and the panel
suggest sizes from what the runs used, and you apply them ([`docs/sizing.md`](docs/sizing.md)).
- **Model allow lists.** A trigger's `run.models`, or `PI_ALLOWED_MODELS` for the whole deployment, names the
models a job may use. The job is stopped before it calls any other model
([`docs/triggers.md`](docs/triggers.md)).
Expand Down Expand Up @@ -288,7 +292,8 @@ and the history, where `r` reverts to an earlier row:
### More than one machine

Several machines can share one queue, one budget and one panel once each worker has a name
(`PI_WORKER_NAME`). Do not share the sandbox directory between them
(`PI_WORKER_NAME`). Each machine keeps its own memory and CPU budget for its jobs, and a job on the shared queue waits
for a machine its size fits on ([`docs/sizing.md`](docs/sizing.md)). Do not share the sandbox directory between them
([`docs/multi-host.md`](docs/multi-host.md) says why).

## Flows: the custom prompt a trigger runs
Expand Down
5 changes: 5 additions & 0 deletions admin/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -161,6 +161,11 @@ One command puts a live terminal view over the whole deployment:
no switch: one job per folder at a time (it lives in the worker process, and one worker per container
daemon is the supported shape), because two agents editing one working tree race each other with no
gate and no undo.
- **Job sizes.** A project row can set the memory and CPUs of its jobs, and how much of a machine they may hold
(`hostShare`) or keep room for (`minJobs`). The PROJECTS view (`j`) shows each project's size, the p95 of its
runs' memory peaks and cores used, a suggested size and the exact `dispatch_limit_edit` call that applies it.
`dispatch_limit_add` and `dispatch_limit_edit` take the four fields behind your confirm, and nothing applies a
size by itself ([`docs/sizing.md`](https://github.com/edgehero/pi-dispatch/blob/main/docs/sizing.md)).
- **Dollar windows.** With dollar caps set, the panel shows each window's spend and holds against its cap,
and what the run records settled ([`docs/costs.md`](https://github.com/edgehero/pi-dispatch/blob/main/docs/costs.md)).
- **Projects.** `j` shows each project with its members and this month's spend, and `Enter` on one filters
Expand Down
55 changes: 30 additions & 25 deletions docs/backends.md
Original file line number Diff line number Diff line change
Expand Up @@ -143,31 +143,36 @@ ones adapters get wrong:
--live` reads what these cannot: `pids.max`, `memory.max`, `memory.swap.max`, `cpu.max` and `cpu.weight` inside
a real container, and its `/proc/self/mountinfo`.
- **Every venue on this host runs a job at its size** (issue #596): `--memory` with an equal `--memory-swap` (no
swap beyond it), `--cpu-shares` for its CPU weight, `--shm-size` at most half its memory, one `--cpus` ceiling
on each job (the host's CPU budget, by default the runtime's CPU count minus one core when it has four or more),
and two labels with its size (`pi.dispatch.mem`, `pi.dispatch.cpu`) that doctor holds the
[host budget](multi-host.md#the-host-budget) against. The ceiling bounds a single job; every job container also
runs under the one parent cgroup `pidispatch.slice` (`--cgroup-parent`), whose CPU quota is the CPU budget, so all
jobs together keep the reserve ([the CPU reserve across all jobs](multi-host.md#the-cpu-reserve-across-all-jobs)).
The egress proxy and Valkey never run under it. Who sets that quota depends on the venue: the worker itself on
rootless Podman (`systemctl --user`) and on Docker Desktop (a short helper container that writes `cpu.max`, written
again after a Docker Desktop restart), the operator once as root on Docker with systemd and on rootful Podman
(doctor prints the command). A quota that is not in place never stops a job: doctor warns "no host CPU reserve
across jobs" and the worker logs `cpu_reserve_fail_open`. There is no memory limit across jobs (the kernel would
kill the largest job, not the one that grew); the budget's admission bounds memory. Both builders emit the same
flags. A venue's `info` gives the budget's `auto` its memory
and CPU count; a venue that gives neither leaves the budget unknown, and the worker then holds no job back on it
and says so (`host_budget_unknown`). A container whose stop fails keeps its room in the budget until the venue's own
`ps -a` no longer lists it or lists it as `exited`, `dead` or `stopped` (one still `created` is removed first), so a
venue's container names and its `{{.State}}` words must stay exact. When the worker starts, it lists the venue's job
containers left after its reaper (`ps -a` with `{{.State}}` and the two size labels) and counts each until it is
gone; while that listing fails, it admits no job on that venue (the other venues' jobs still run). A venue whose
binary is not installed, or whose job user is refused, holds no container and counts as none. The size is the job's
project row's, else `PI_JOB_MEMORY` and `PI_JOB_CPUS` ([job sizes](scoped-limits.md#job-sizes-version-3)). Where
Docker reports `SwapLimit` or `CPUShares` false, it drops that flag: doctor warns, and the worker logs
`size_bound_unenforced` for each job. Where the runtime gives no CPU count, jobs run with no `--cpus` and doctor
warns. On Docker Desktop, after you give the VM fewer CPUs the first job loses one attempt (refunded and retried)
and the worker logs `cpu_ceiling_stale`; the next pickup reads the new count.
swap beyond it), `--cpu-shares` for its CPU weight, `--shm-size` at most half its memory, one `--cpus` ceiling on
each job (the host's CPU budget, by default the runtime's CPU count minus one core when it has four or more), and
two labels with its size (`pi.dispatch.mem`, `pi.dispatch.cpu`) that doctor holds the [host
budget](multi-host.md#the-host-budget) against. The ceiling bounds a single job; every job container also runs
under the one parent cgroup `pidispatch.slice` (`--cgroup-parent`), whose CPU quota is the CPU budget, so all jobs
together keep the reserve ([the CPU reserve across all jobs](multi-host.md#the-cpu-reserve-across-all-jobs)). The
egress proxy and Valkey never run under it. Who sets that quota depends on the venue: the worker itself on
rootless Podman (`systemctl --user`) and on Docker's `cgroupfs` driver, Docker Desktop included (a short helper
container that writes `cpu.max`, written again after a Docker Desktop restart), the operator once as root on
Docker with systemd and on rootful Podman, and as the daemon's account on rootless Docker (doctor prints the
command). A quota that is not in place never stops a job: doctor warns "no host CPU reserve across jobs" and the
worker logs `cpu_reserve_fail_open`. Doctor reads the quota back on rootless Podman, on the `cgroupfs` driver and
on Docker with systemd on this host; on rootful Podman, rootless Docker and a daemon on another machine nothing
can read it from here, so that warning stays even after the command ran. There is no memory limit across jobs (the
kernel would kill the largest job, not the one that grew); the budget's admission bounds memory. Both builders
emit the same flags. A venue's `info` gives the budget's `auto` its memory and CPU count; a venue that gives
neither leaves the budget unknown, and the worker then holds no job back on it and says so
(`host_budget_unknown`). A container whose stop fails keeps its room in the budget until the venue's own `ps -a`
no longer lists it or lists it as `exited`, `dead` or `stopped` (one still `created` is removed first), so a
venue's container names and its `{{.State}}` words must stay exact. When the worker starts, it lists the venue's
job containers left after its reaper (`ps -a` with `{{.State}}` and the two size labels) and counts each until it
is gone; while that listing fails, it admits no job on that venue (the other venues' jobs still run). A venue
whose binary is not installed, or whose job user is refused, holds no container and counts as none. The size is
the job's project row's, else `PI_JOB_MEMORY` and `PI_JOB_CPUS` ([job
sizes](scoped-limits.md#job-sizes-version-3)). Where Docker reports `SwapLimit` or `CPUShares` false, it drops
that flag: doctor warns, and the worker logs `size_bound_unenforced` for each job. Where the runtime gives no CPU
count, jobs run with no `--cpus` and doctor warns. On Docker Desktop, after you give the VM fewer CPUs the first
job loses one attempt (refunded and retried) and the worker logs `cpu_ceiling_stale`; the next pickup reads the
new count. How to choose sizes and a budget, and what is not isolated (disk I/O, disk space, network bandwidth),
is in [sizing jobs](sizing.md).
- **`local`'s `nonRoot` and `localFolders` depend on which uid the job runs as** (issue #341). On macOS, Windows
and Docker Desktop the image's own `USER` runs. On a daemon that enforces bind-mount ownership (native Linux
Docker, rootful Podman) the worker runs the job as its own uid with `--user` and `HOME=/home/pi`, because
Expand Down
3 changes: 2 additions & 1 deletion docs/insights.md
Original file line number Diff line number Diff line change
Expand Up @@ -146,7 +146,8 @@ the memory part of it (below).
retries. A job where only a child process was killed keeps its own outcome, and `oomKills` above 0 shows it.
- Each job runs at its size: its project's `memory` and `cpus`, else `PI_JOB_MEMORY` and `PI_JOB_CPUS` (default
4 GB and 2 CPUs). Every record says which in `size` (`memMiB`, `cpuCenti` in hundredths of a CPU, and `source`:
`project`, `env` or `default`). Compare `memPeak` with it to choose a size ([job sizes](scoped-limits.md#job-sizes-version-3)).
`project`, `env` or `default`). Compare `memPeak` with it to choose a size ([job sizes](scoped-limits.md#job-sizes-version-3),
and [sizing jobs](sizing.md) for the whole path).
- Disk I/O, disk space and network are not measured or isolated.

### The job sizes section
Expand Down
16 changes: 11 additions & 5 deletions docs/multi-host.md
Original file line number Diff line number Diff line change
Expand Up @@ -175,6 +175,8 @@ for the whole deployment, not per host.

## The host budget

How to choose sizes and read what doctor says about them, with a worked example, is in [sizing jobs](sizing.md).

`PI_CONCURRENCY` counts jobs, and a count cannot tell a 20g job from a 2g one. So each worker also keeps a
**budget** of memory and CPU for its jobs, and starts a job only when its size (its project's `memory` and `cpus`,
see [job sizes](scoped-limits.md#job-sizes-version-3)) fits beside the sizes of the jobs already running on that
Expand Down Expand Up @@ -258,11 +260,15 @@ Who sets the quota depends on the container runtime:
sudo systemctl set-property pidispatch.slice CPUQuota=300%
```

`pi-dispatch doctor` says per runtime whether the quota is in place, and warns "no host CPU reserve across jobs" with
the fix when it is not. Jobs still run then, inside the parent but without the quota (the worker logs
`cpu_reserve_fail_open` with the reason). With `PI_HOST_CPU_BUDGET=off` no quota is set, and the worker removes one
it set before. Where rootless Podman uses the cgroupfs manager instead of systemd, jobs run without the parent and
each job's CPU weight is capped at the default: the proxy and Valkey then get a fair share of the CPU, not a reserve.
On rootless Podman, Docker Desktop and Docker's cgroupfs driver, and Docker with systemd on this host, `pi-dispatch
doctor` reads the quota back and warns "no host CPU reserve across jobs" with the command when it is missing or
differs from the budget. On rootful Podman, rootless Docker and a Docker daemon on another machine nothing can read
it from here, so doctor's warning stays even after you ran the command, and the worker logs `cpu_reserve_fail_open`
with the status `unmanaged`. On a cgroup v1 host no quota is kept. Jobs run in every case, inside the parent. With
`PI_HOST_CPU_BUDGET=off` the worker clears the quota only where it sets it itself (rootless Podman, Docker's
cgroupfs driver); one set with `sudo` stays until you clear it with `sudo systemctl set-property pidispatch.slice
CPUQuota=`. Where rootless Podman uses the cgroupfs manager instead of systemd, jobs run without the parent and each
job's CPU weight is capped at the default: the proxy and Valkey then get a fair share of the CPU, not a reserve.

There is no memory limit across all jobs, on purpose: when a group of containers runs out of memory together, the
kernel kills the largest job in the group, not the one that grew (measured). The budget's admission keeps the jobs'
Expand Down
2 changes: 2 additions & 0 deletions docs/podman.md
Original file line number Diff line number Diff line change
Expand Up @@ -1251,6 +1251,8 @@ own folder is never relabelled and needs the `semanage fcontext` label from the

### The CPU reserve on this venue

How to choose sizes and a budget is in [sizing jobs](sizing.md); this section is what differs on this venue.

Every job gets the size its project sets (see [job sizes](scoped-limits.md#job-sizes-version-3)) and a `--cpus`
ceiling of the host's CPU budget ([the host budget](multi-host.md#the-host-budget)), which on this venue also takes
a `cpu.max` or `memory.max` set on the account's systemd user service into account. The ceiling bounds one job; what
Expand Down
2 changes: 1 addition & 1 deletion docs/projects.md
Original file line number Diff line number Diff line change
Expand Up @@ -125,7 +125,7 @@ Caps live in [`scoped-limits.json`](scoped-limits.md#project-rows), never in thi
- A project row needs `"version": 2` in `scoped-limits.json`, even with counts only. The panel and the tools write it.
- The same row can set the size of every member's job container: `"memory": "8g"` and `"cpus": 4` (file version
3). Without it a job gets `PI_JOB_MEMORY` and `PI_JOB_CPUS` (default `4g` and `2`). See
[job sizes](scoped-limits.md#job-sizes-version-3).
[job sizes](scoped-limits.md#job-sizes-version-3), and [sizing jobs](sizing.md) for how to choose one.

The row's id must be a project here. A row naming a missing project stops the worker from starting, and a live
edit that would leave one is kept out (the worker logs `scoped_limits_reload_invalid` or `projects_reload_invalid`,
Expand Down
3 changes: 2 additions & 1 deletion docs/scoped-limits.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,7 +210,8 @@ Things to know before you add a model row:
### Job sizes (version 3)

A `project:<id>` row can also say how big each of the project's job containers is. Heavy projects get more
memory, light ones less.
memory, light ones less. This section is the reference for the four fields; [sizing jobs](sizing.md) walks through
measuring, setting sizes and the host budget, with a worked example.

```json
{
Expand Down
Loading
Loading