Skip to content

docs(sizing): add the operator guide to job sizes and the host budget (#596) - #602

Merged
edgehero merged 2 commits into
mainfrom
docs/596-sizing
Oct 7, 2026
Merged

edgehero merged 2 commits into
mainfrom
docs/596-sizing

Conversation

@edgehero

@edgehero edgehero commented Oct 7, 2026

Copy link
Copy Markdown
Owner

Phase 4 of #596: one operator guide, docs/sizing.md, that walks through sizing jobs in the order an operator meets it: measure first, set a size per project, the host budget, fairness (minJobs and hostShare), the CPU reserve across all jobs, and reading the suggestions. It states plainly what is isolated (memory, CPU, processes) and what is not (disk I/O, disk space, network bandwidth), and that page cache counts toward a job's memory peak.

A worked example sizes a 64 GB, 16 CPU host for three projects (20g/4, 8g/2, 2g/1) with the exact files, doctor's exact lines, and what happens under a flood of each kind. A new test reads the guide's own files through the worker's loaders and replays both floods through the real budget, so the example cannot drift from the code.

The README, the admin README and the reference docs (scoped-limits, multi-host, projects, insights, backends, podman) link to the guide. docs/backends.md also gains the two venues its reserve list missed (Docker's cgroupfs driver and rootless Docker). No spec or behaviour change.

Full suite: 7268 tests, 0 failures.

Refs #596

…#596)

Phases 0 to 3 documented each piece where it lives (job sizes in
scoped-limits, the budget and the CPU reserve in multi-host, the
suggestions and the resources block in insights), but nothing walked
an operator from "every job is 4g" to sizes that fit their projects.
docs/sizing.md is that path, in order: measure first, set a size per
project, the host budget (auto, reserves, off, PI_CONCURRENCY as the
count, waiting versus refused, single host versus a fleet), fairness
(minJobs, hostShare, the room kept for the oldest waiting job), the CPU
reserve across all jobs per runtime, and reading the suggestions. It
states plainly that only memory and CPU are isolated: disk I/O, disk
space and network bandwidth are shared, and the page cache counts
toward a job's memory peak.

A worked example (64 GB, 16 CPUs, heavy 20g/4, medium 8g/2, light 2g/1)
gives the exact projects file, limits file and env, doctor's lines, two
floods and an out of memory suggestion. Those restate derivable
sources, so worker/test/sizing-doc.test.mjs bolts them: it reads the
page's own JSON and env blocks, feeds them to the worker's loaders and
doctor's functions, requires the printed lines verbatim, drives both
floods through makeHostBudget, and holds the CPU reserve table to
reservePlan. Six mutations of the page, each killed.

The READMEs gain a feature bullet (admin: the PROJECTS view and the
four row fields) and the fleet sentence; scoped-limits, multi-host,
projects, insights, podman and backends link to the guide instead of
repeating it. backends.md now names Docker's cgroupfs driver and
rootless Docker in who sets the quota, as reservePlan decides.

Specs UNCHANGED, checked: REQ-SCOPED-LIMITS, DES-HOST-BUDGET,
DES-SIZE-SUGGESTIONS, INT-SCOPED-LIMITS-FILE-CONTRACT and
INT-CONTAINER-RUNTIME-CONTRACT agree with every claim the guide makes.

Refs #596

Signed-off-by: Rob Boerman <robboerman@live.nl>
…o act on a suggestion (#596)

Doctor does not read the quota of pidispatch.slice back on every runtime.
reservePlan gives no method for rootful or remote Podman, rootless Docker,
a systemd Docker daemon on another machine and cgroup v1, so there the
warning stays even after the operator ran the command and the worker logs
cpu_reserve_fail_open with the status unmanaged. sizing.md, multi-host.md
and backends.md now say which runtimes are read back (rootless Podman, the
cgroupfs driver including Docker Desktop, Docker with systemd on this host)
and which are not, and the table gains that column.

With PI_HOST_CPU_BUDGET=off the worker clears only a quota it sets itself;
one set with sudo stays until cleared, and the page quotes doctor's warning.

sizing.md also says how to apply a call (ask pi with the admin extension to
run it; the panel's limit dialogs keep a row's size), how to turn the budget
off and set a reserve, that the four budget settings are read at worker
start, what the sudo command does and how to remove it, that Docker
Desktop's budget is the VM's and the helper needs the job image, that
suggestions read this machine's run records only, and that /dev/shm counts
toward memory. Jargon is glossed on first use. Wording fixes on the OOM
raise, which runs count, when a call is offered, and in README.md and both
.env.example copies (the CPU ceiling is the budget when one is in force).

sizing-doc.test.mjs now pins the prose numbers: the flood sums from the
budget's entries, both "Seven start" counts, the re-check cadences, the
auto reserve, the shm and process caps, the run counts, the suggestion
factors (checked against suggestSize at their boundary), the four CPU
budget sentence, the standalone root commands and the table's read-back
column. Lines are printed with doctor's own render.

Specs UNCHANGED, checked: DES-HOST-BUDGET and DES-SIZE-SUGGESTIONS already
describe this behaviour; only the operator docs were wrong.

Refs #596

Signed-off-by: Rob Boerman <robboerman@live.nl>
@edgehero
edgehero merged commit dc2455e into main Oct 7, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant