Repository navigation
Size each job container per project, within a budget per host, so no job can starve the others #596
Copy link
Copy link
Closed
Description
Activity
Plan
The design was settled on 2026-10-06. The work comes in five phases, each its own pull request.
Decided
- Measure first. The first release only records each job's real peak memory and CPU, and handles out-of-memory kills, while every job keeps today's 4 GB and 2 CPUs. Sizes are chosen afterwards, from real numbers.
- CPU: a fair share plus a host ceiling.
- Each job gets a CPU weight in proportion to its size. Under contention, every job therefore gets at least its size.
- A job running alone may use idle cores, up to the host budget, but never the reserve kept for the worker and the system.
- This replaces a hard cap at the job's size, which leaves idle cores unused.
- The budget defaults to
auto. Each host's budget is what the container runtime reports, minus a reserve, and it always fits at least one default job. The release notes will say that small machines may run fewer jobs at once than they do today. - Fairness comes from a soft minimum per project.
minJobson a project row keeps room for N of that project's jobs whenever one is waiting.- An optional
hostSharecaps the share of the host a project may hold. - When only one project is busy, no capacity sits idle. A waiting heavy or light project starts within one job's duration.
Defaults, unless changed
- Sizes and where they live:
- Sizes are raw numbers (
memory,cpus), not classes. - They live on project rows in
scoped-limits.jsonversion 3, withPI_JOB_MEMORYandPI_JOB_CPUSas the deployment default. - There is no size per trigger.
- Sizes are raw numbers (
- Container limits:
- No swap beyond memory:
--memory-swapequals--memory. --pids-limitstays at 512.- Shared memory is
min(1g, memory / 2). - Floors: 512m memory and 0.25 CPUs.
- No swap beyond memory:
- Out-of-memory kills:
- A confirmed one is recorded
oom-killedand is not retried. - An unconfirmed exit 137 is retried as today.
- The comment is generic. The project and size go to the log, the record and doctor.
- A confirmed one is recorded
- A container whose stop failed keeps its budget hold until the runtime confirms it is gone.
- The reserve: 10% of memory, clamped to 1 to 4 GB, and 1 CPU on hosts with 4 or more.
- Out of scope: disk I/O, disk space and network are not isolated. The docs say so.
- Size suggestions are shown, never applied automatically.
Phases
Phase What 0 Measure. The runner reads its cgroup at exit ( memory.peak,memory.events,cpu.stat,pids.peak) into the signed exit line, and the run record gainsresources. Docker'soomevent stream feeds theoom-killedhandling. A lab measures what each venue exposes, including how rootless Podman can report an out-of-memory kill.1 Sizes and hard limits. PI_JOB_MEMORYandPI_JOB_CPUS; projectmemory,cpus,hostShareandminJobs(scoped limits v3); no swap; CPU shares; the full list of forbidden extra flags. The size goes in the record, and the sandbox, probes and canary use the real size.2 Host budget. autodetection with the reserve, and a budget gate that runs last. One release path holds every job's resources, and orphans are kept on a failed stop. Holds for the oldest waiter and for each project below itsminJobs; thehostSharecap; never-fits refusals per host and per fleet; the fleet view in doctor.3 Suggestions. Sizes per project from recorded peaks, shown in doctor, the panel and insights. 4 Docs and release. - added 14 commits that reference this issue
on Oct 6, 2026 9 remaining items
- added 13 commits that reference this issue
on Oct 7, 2026 Done in pi-dispatch 4.0.0 (https://github.com/edgehero/pi-dispatch/releases/tag/v4.0.0).
What shipped, by phase
- Measure first (feat(runner,worker): record each run's memory and cpu use, and stop retrying a job killed for memory (#596) #597). Every run records what it used (
resources: peak memory and swap, out of memory kills, memory pressure, CPU time and throttling, peak processes), signed by the runner. A small supervisor in the job image reports a kill the job cannot report itself, so a confirmed out of memory kill is recordedoom-killedand not retried or paid twice. Review found and closed a hole open since 3.0.0: a job could open Node's debugger on the runner and read the key that signs its result. - A size per project (feat(worker,admin): size each job container per project, with no swap beyond its memory (#596) #598).
memoryandcpuson aproject:row (scoped limits version 3),PI_JOB_MEMORYandPI_JOB_CPUSas defaults, no swap beyond memory, CPU weight from the size. - A budget per machine (feat(worker,admin): hold each host's jobs inside a memory, CPU and count budget, with one CPU reserve across all jobs (#596) #600).
PI_HOST_MEMORY_BUDGETandPI_HOST_CPU_BUDGET(autoby default, with a reserve for the machine),PI_CONCURRENCYas the job count, holds so a big job is not starved,minJobsandhostShareenforced, sizes that can never fit refused before any spend, and one CPU limit across all jobs throughpidispatch.sliceso a core stays free for the machine where the runtime allows it. - Suggestions (feat(worker,admin): suggest a job size per project from its runs, capped at the host, never applied automatically (#596) #601).
doctor, the panel, the limit edit preview and insights suggest a size per project from the measured runs, capped at what the machine offers, never applied automatically. Memory is raised only after a confirmed kill; CPU is only lowered. - Docs (docs(sizing): add the operator guide to job sizes and the host budget (#596) #602).
docs/sizing.md, with a worked example for a 64 GB machine shared by a heavy, a medium and a light project.
Evidence
- Measured on five setups (Docker Desktop, Docker 29, rootless Podman 4.9, Podman 5.8 rootful and rootless): each size read back exactly from the container's cgroup (memory, swap 0, CPU limit and weight, processes, shared memory); a job past its memory stopped alone while its neighbour ran on; two busy jobs split CPU in size order; with the CPU limit across jobs, busy jobs together stayed at the budget while a busy program outside kept its core.
- Every phase went through review rounds (correctness, an adversarial reviewer, runs on the five setups), with the review scenarios kept as scripts and re-run after each fix. Every new test pin was checked by deliberately breaking the code.
- The full suite passed on every merge (7272 tests at release), and all six required checks were green.
Known limits, stated in the docs
- Disk I/O, disk space and network bandwidth are not isolated between jobs.
- On Linux with Docker's systemd driver or rootful Podman, the CPU limit across jobs needs one
sudocommand, whichdoctorprints; on rootful Podman and remote daemonsdoctorcannot read it back. - Suggestions read this machine's run records only, unless
PI_LOGS_DIRis shared. - The panel's limit dialogs keep a project's size but do not prompt for one; a suggested size is applied by asking pi to run the call (it asks you to confirm).
The machine-level view of how busy each host is, and was, is #599.
- Measure first (feat(runner,worker): record each run's memory and cpu use, and stop retrying a job killed for memory (#596) #597). Every run records what it used (
Metadata
Metadata
Assignees
Labels
No labels
Goal
Every job container gets a fixed size today: 4 GB memory, 2 CPUs and 512 processes (
worker/src/container-spec.mjs:151-152,worker/src/docker-run.mjs:42). Nothing can change it, and nothing checks that the jobs running on a host fit that host.The goal:
20 GB per job on a 64 GB machine is one example. It must work for any size and any host.
What exists to build on
PI_CONCURRENCYcaps jobs per host. Scoped limitsconcurrentcaps jobs per repo, folder or project (project:<id>rows), and an excess job is deferred, never dropped (worker/src/index.mjs:583-685). The folder mutex keeps one job per local folder.effectiveJobOf(index.mjs:167) already resolves per-job facts. The job image already travels the same way intomakeRunContainer→buildArgs(run-container.mjs:33,:163-190).buildDockerRunArgsandbuildPodmanRunArgspass their options straight tocontainerSpec, so memory and cpus need only be passed along.DAEMON_APPLIES_BOUNDSandPODMAN_BOUNDS_DELEGATED(backends.mjs:348-376,backend-podman.mjs:254-300) already tell whether the runtime really applies pids and memory bounds, and the isolation word andPI_BACKEND_FLOORact on that.scoped-limits.mjs) are where per-project settings already live, with admin tools, the panel and doctor around them.Design
1. A size per project, and a deployment default
PI_JOB_MEMORYandPI_JOB_CPUSset it, and both stay at4gand2when unset, so nothing changes for existing deployments.512m,8gor20g, and CPUs as a positive number with at most two decimals..env.example.project:<id>row inscoped-limits.jsonmay carrymemoryandcpus.index.mjs:594-602), and is never stored in the job data. Cron templates and jobs already queued or deferred therefore always use the current size.run.networkwas refused for (design.md:5474-5476,interfaces.md:7271). The operator sizes projects; a trigger never does.2. Hard limits, so no container takes another's share
Every job container gets:
--memory=<size>together with--memory-swap=<same size>.--memory-swapis only refused as an extra flag, so Docker's default can give a container swap beyond its memory.--cpus=<n>stays a hard ceiling.--cpu-sharesproportional to the job's CPUs gives each job its fair share when every running job wants CPU at once.--pids-limitstays, unchanged.PI_BACKEND_FLOORsuch a host refuses jobs before they spend.--cpuson every job, as today. Without the cpu controller delegated, no job starts there (exit 126,docs/podman.md:628-632). doctor already names this and stays the place to fix it.3. A budget per host, so the jobs on a host always fit
PI_HOST_MEMORY_BUDGETandPI_HOST_CPU_BUDGET, defaultauto.autois what the container runtime reports for this host (docker info/podman info: MemTotal and NCPU) minus a reserve for the worker, Valkey and the system (PI_HOST_RESERVE_MEMORY,PI_HOST_RESERVE_CPUS, with sensible defaults).moveToDelayedpath, and its earlier holds are given back in the same drain.PI_CONCURRENCYstays as a plain upper bound on the number of jobs, and projectconcurrentkeeps working as it does.job-size-exceeds-host). It never waits forever.hostShare: 0.5). Its running jobs together then never hold more than that share, so one project with a long queue cannot fill the machine while another project's job waits.4. Measure, then tune: the balance per project
The run record gains its size (memory and CPUs) and the peak memory and CPU use the job really had. The fields go at the tail, so older records stay valid.
Out of memory becomes visible and stops costing double.
processor.mjs:1340-1344,:1460).oom-killed.doctor reports, per project:
It also warns when the project sizes and shares cannot all fit on this host.
The panel and the insights page show the host budget in use and free, and each project's size against its peaks. The admin tools (
dispatch_limit_addanddispatch_limit_edit) learnmemory,cpusandhostShare, confirm-gated as today.5. Everything that assumes 4 GB today moves with it
sandbox.mjs:187-210).live-probes.mjs:115-149,:258-294).doctor.mjs:6180-6206).Phases
Each phase is its own PR, with specs, tests, mutation checks and review rounds, as in the last rounds.
PI_JOB_MEMORYandPI_JOB_CPUS, projectmemoryandcpus(scoped limits v3), no swap beyond memory, CPU shares, the size in the run record, and the sandbox, probes and canary using the real size.autodetection and the reserve, the wait-longest hold,hostShare, and the never-fits refusal.oom-killedhandling, doctor suggestions, and the panel, insights and admin tools.docs/scoped-limits.md,docs/projects.md,docs/multi-host.md,docs/podman.md, and the release notes.Specs to amend
INT-CONTAINER-RUNTIME-CONTRACT(interfaces.md:1378-1392,:1954-1967).INT-SANDBOX-CONTRACT.INT-LIVE-PROBE-CONTRACT.CONST-ISOLATION-CONTAINER-PER-JOB(constitution.md:291).DES-CONCURRENCY-3, and OQ-002, which says the memory per job was never measured. The recorded peaks finally measure it.INT-SCOPED-LIMITS-FILE-CONTRACT,REQ-SCOPED-LIMITSandDES-SCOPED-LIMITS-AND-FOLDER-MUTEX.INT-RUN-HISTORY-FILE-CONTRACT.DES-PODMAN-NATIVE-ROOTLESS-BACKEND.Open questions
hostShareplus the wait-longest hold enough?Acceptance
oom-killed, and the job is not retried.