Skip to content

Size each job container per project, within a budget per host, so no job can starve the others #596

Description

@edgehero

Goal

Every job container gets a fixed size today: 4 GB memory, 2 CPUs and 512 processes (worker/src/container-spec.mjs:151-152, worker/src/docker-run.mjs:42). Nothing can change it, and nothing checks that the jobs running on a host fit that host.

The goal:

  1. Each project gets the compute it needs. One project may need a lot of memory and CPU, another very little.
  2. Each host is used well. The same configuration should give a sensible balance on a 16 GB laptop and on a 128 GB server.
  3. No container can take compute that another running container was promised. A runaway or heavy job hits its own limit. It never slows the host or starves the other jobs.

20 GB per job on a 64 GB machine is one example. It must work for any size and any host.

What exists to build on

  • Concurrency. PI_CONCURRENCY caps jobs per host. Scoped limits concurrent caps jobs per repo, folder or project (project:<id> rows), and an excess job is deferred, never dropped (worker/src/index.mjs:583-685). The folder mutex keeps one job per local folder.
  • The per-job seam. effectiveJobOf (index.mjs:167) already resolves per-job facts. The job image already travels the same way into makeRunContainer → buildArgs (run-container.mjs:33, :163-190). buildDockerRunArgs and buildPodmanRunArgs pass their options straight to containerSpec, so memory and cpus need only be passed along.
  • Runtime observations. DAEMON_APPLIES_BOUNDS and PODMAN_BOUNDS_DELEGATED (backends.mjs:348-376, backend-podman.mjs:254-300) already tell whether the runtime really applies pids and memory bounds, and the isolation word and PI_BACKEND_FLOOR act on that.
  • The projects file and scoped limits v2 (scoped-limits.mjs) are where per-project settings already live, with admin tools, the panel and doctor around them.

Design

1. A size per project, and a deployment default

  • Deployment default. PI_JOB_MEMORY and PI_JOB_CPUS set it, and both stay at 4g and 2 when unset, so nothing changes for existing deployments.
    • Both are validated at load: a memory size like 512m, 8g or 20g, and CPUs as a positive number with at most two decimals.
    • Both go in .env.example.
  • Per project. A project:<id> row in scoped-limits.json may carry memory and cpus.
    • This needs scoped limits version 3. Today unknown fields are dropped silently, so an older worker must refuse the file instead of running the project at the default size.
    • A repo or folder row may carry a size too, if that turns out useful. Project rows come first.
  • When the size is chosen. It is resolved at pickup, from the same limits snapshot the project gate reads (index.mjs:594-602), and is never stored in the job data. Cron templates and jobs already queued or deferred therefore always use the current size.
  • No per-trigger size. A trigger raising its own container size is a per-trigger relaxation of a safety bound, the shape run.network was refused for (design.md:5474-5476, interfaces.md:7271). The operator sizes projects; a trigger never does.
    • Not having one also avoids a new receiver field and its skew rules.

2. Hard limits, so no container takes another's share

Every job container gets:

  • Memory: a hard cap with no extra swap. --memory=<size> together with --memory-swap=<same size>.
    • Today --memory-swap is only refused as an extra flag, so Docker's default can give a container swap beyond its memory.
    • It becomes a spec field set by the worker, never by a trigger or extra flags.
  • CPU.
    • --cpus=<n> stays a hard ceiling.
    • --cpu-shares proportional to the job's CPUs gives each job its fair share when every running job wants CPU at once.
  • Processes. --pids-limit stays, unchanged.
  • Disk I/O. I/O weight, so one job cannot starve the disk. Verify it is applied on Docker and on rootful and rootless Podman; where it is not, name it, never assume it.
  • Isolation checks.
    • The existing runtime observations extend to CPU wherever the runtime can show it applies a CPU limit.
    • doctor names any limit that is not enforced, and with PI_BACKEND_FLOOR such a host refuses jobs before they spend.
    • Rootless Podman keeps --cpus on every job, as today. Without the cpu controller delegated, no job starts there (exit 126, docs/podman.md:628-632). doctor already names this and stays the place to fix it.

3. A budget per host, so the jobs on a host always fit

  • Settings. PI_HOST_MEMORY_BUDGET and PI_HOST_CPU_BUDGET, default auto.
    • auto is what the container runtime reports for this host (docker info / podman info: MemTotal and NCPU) minus a reserve for the worker, Valkey and the system (PI_HOST_RESERVE_MEMORY, PI_HOST_RESERVE_CPUS, with sensible defaults).
    • The runtime's own numbers also cover Docker Desktop, whose VM has less than the machine.
  • A new pickup gate. It sits beside the existing scope and project gates:
    • A job starts only when its memory and CPUs fit what the running jobs on this host leave free. Otherwise it is deferred like a busy scope, through the same moveToDelayed path, and its earlier holds are given back in the same drain.
    • The in-process count is correct for the same reason as today: one worker per container daemon (DES-CONCURRENCY-3).
  • Guaranteed promises. The sizes of running jobs never add up to more than the budget, so every running job always gets its full size.
  • Limits that still apply. PI_CONCURRENCY stays as a plain upper bound on the number of jobs, and project concurrent keeps working as it does.
  • A job that can never fit. If its size exceeds the whole host budget, it is refused before it spends, with a reason that names the sizes (for example job-size-exceeds-host). It never waits forever.
  • No starvation. When a big job has waited longest, the budget it needs is held for it. Smaller jobs keep starting only while they fit beside that hold. Without this, a stream of small jobs could keep a 20 GB job waiting forever.
  • Fair use across projects. A project row may set a share of the host budget (for example hostShare: 0.5). Its running jobs together then never hold more than that share, so one project with a long queue cannot fill the machine while another project's job waits.
    • A guaranteed minimum per project, like the budget split's floors, is an open question (see below).
  • Replicas. Each replica counts toward the budget.

4. Measure, then tune: the balance per project

  • The run record gains its size (memory and CPUs) and the peak memory and CPU use the job really had. The fields go at the tail, so older records stay valid.

  • Out of memory becomes visible and stops costing double.

    • Today an out-of-memory kill is an unclassified exit 137: it is retried, paid twice and recorded with no reason (processor.mjs:1340-1344, :1460).
    • The worker reads the runtime's OOM-killed state for that container and records the reason oom-killed.
    • It does not retry, because the same size would fail the same way.
    • The comment says which project's size to raise.
  • doctor reports, per project:

    • the size;
    • the recent peaks;
    • a suggestion: "shop jobs peak at 6 GB of their 20 GB; 8 GB would do" or "platform jobs were killed for memory twice this week; raise its memory".

    It also warns when the project sizes and shares cannot all fit on this host.

  • The panel and the insights page show the host budget in use and free, and each project's size against its peaks. The admin tools (dispatch_limit_add and dispatch_limit_edit) learn memory, cpus and hostShare, confirm-gated as today.

5. Everything that assumes 4 GB today moves with it

  • The sandbox reopens a job at that job's recorded size (sandbox.mjs:187-210).
  • The live probes check what jobs really get, not the old default (live-probes.mjs:115-149, :258-294).
  • The doctor egress canary gets the real size too (doctor.mjs:6180-6206).
  • Fleet. Project sizes are fleet-wide through the scoped limits file and its fingerprint, and doctor names a host that disagrees. The host budget is per host by design.
    • A host too small for a project's size refuses that project's local jobs with a named reason. A forge job is left for a host that fits.

Phases

Each phase is its own PR, with specs, tests, mutation checks and review rounds, as in the last rounds.

  1. Size and hard limits. PI_JOB_MEMORY and PI_JOB_CPUS, project memory and cpus (scoped limits v3), no swap beyond memory, CPU shares, the size in the run record, and the sandbox, probes and canary using the real size.
  2. Host budget. The budget gate with auto detection and the reserve, the wait-longest hold, hostShare, and the never-fits refusal.
  3. Measure and suggest. Peak use in the record, oom-killed handling, doctor suggestions, and the panel, insights and admin tools.
  4. Docs and release. README, docs/scoped-limits.md, docs/projects.md, docs/multi-host.md, docs/podman.md, and the release notes.

Specs to amend

  • INT-CONTAINER-RUNTIME-CONTRACT (interfaces.md:1378-1392, :1954-1967).
  • INT-SANDBOX-CONTRACT.
  • INT-LIVE-PROBE-CONTRACT.
  • CONST-ISOLATION-CONTAINER-PER-JOB (constitution.md:291).
  • DES-CONCURRENCY-3, and OQ-002, which says the memory per job was never measured. The recorded peaks finally measure it.
  • INT-SCOPED-LIMITS-FILE-CONTRACT, REQ-SCOPED-LIMITS and DES-SCOPED-LIMITS-AND-FOLDER-MUTEX.
  • INT-RUN-HISTORY-FILE-CONTRACT.
  • DES-PODMAN-NATIVE-ROOTLESS-BACKEND.

Open questions

  • Guaranteed minimum per project. Should a project be able to reserve part of the host budget even when it has no jobs running, like the budget split's floors? Or is hostShare plus the wait-longest hold enough?
  • How to read peak use. Choose between the cgroup's peak counters before the container is removed, and runtime stats sampled while the job runs. Measure both on Docker and on rootless Podman.
  • Disk I/O weight under rootless Podman may not be enforceable. Measure it before promising it.
  • Defaults for the reserve. For example 2 GB and 1 CPU, or a fraction of the host.

Acceptance

  • A deployment with no new settings runs every job exactly as today: 4 GB, 2 CPUs, 512 processes.
  • On a 64 GB host with three projects sized 20 GB, 8 GB and 2 GB, the running jobs never hold more than the budget. A job that does not fit waits and starts as soon as it fits. A big waiting job is not starved by small ones.
  • A job that tries to use more memory than its size is stopped on its own, while the other running jobs keep their memory and CPU. The record says oom-killed, and the job is not retried.
  • A job whose size can never fit the host is refused before it spends, with a reason naming both sizes.
  • doctor names any limit the runtime does not enforce, and suggests a better size per project from real peaks.

Activity

  1. edgehero commented on Oct 6, 2026

    @edgehero
    OwnerAuthor

    Plan

    The design was settled on 2026-10-06. The work comes in five phases, each its own pull request.

    Decided

    1. Measure first. The first release only records each job's real peak memory and CPU, and handles out-of-memory kills, while every job keeps today's 4 GB and 2 CPUs. Sizes are chosen afterwards, from real numbers.
    2. CPU: a fair share plus a host ceiling.
      • Each job gets a CPU weight in proportion to its size. Under contention, every job therefore gets at least its size.
      • A job running alone may use idle cores, up to the host budget, but never the reserve kept for the worker and the system.
      • This replaces a hard cap at the job's size, which leaves idle cores unused.
    3. The budget defaults to auto. Each host's budget is what the container runtime reports, minus a reserve, and it always fits at least one default job. The release notes will say that small machines may run fewer jobs at once than they do today.
    4. Fairness comes from a soft minimum per project.
      • minJobs on a project row keeps room for N of that project's jobs whenever one is waiting.
      • An optional hostShare caps the share of the host a project may hold.
      • When only one project is busy, no capacity sits idle. A waiting heavy or light project starts within one job's duration.

    Defaults, unless changed

    • Sizes and where they live:
      • Sizes are raw numbers (memory, cpus), not classes.
      • They live on project rows in scoped-limits.json version 3, with PI_JOB_MEMORY and PI_JOB_CPUS as the deployment default.
      • There is no size per trigger.
    • Container limits:
      • No swap beyond memory: --memory-swap equals --memory.
      • --pids-limit stays at 512.
      • Shared memory is min(1g, memory / 2).
      • Floors: 512m memory and 0.25 CPUs.
    • Out-of-memory kills:
      • A confirmed one is recorded oom-killed and is not retried.
      • An unconfirmed exit 137 is retried as today.
      • The comment is generic. The project and size go to the log, the record and doctor.
    • A container whose stop failed keeps its budget hold until the runtime confirms it is gone.
    • The reserve: 10% of memory, clamped to 1 to 4 GB, and 1 CPU on hosts with 4 or more.
    • Out of scope: disk I/O, disk space and network are not isolated. The docs say so.
    • Size suggestions are shown, never applied automatically.

    Phases

    Phase What
    0 Measure. The runner reads its cgroup at exit (memory.peak, memory.events, cpu.stat, pids.peak) into the signed exit line, and the run record gains resources. Docker's oom event stream feeds the oom-killed handling. A lab measures what each venue exposes, including how rootless Podman can report an out-of-memory kill.
    1 Sizes and hard limits. PI_JOB_MEMORY and PI_JOB_CPUS; project memory, cpus, hostShare and minJobs (scoped limits v3); no swap; CPU shares; the full list of forbidden extra flags. The size goes in the record, and the sandbox, probes and canary use the real size.
    2 Host budget. auto detection with the reserve, and a budget gate that runs last. One release path holds every job's resources, and orphans are kept on a failed stop. Holds for the oldest waiter and for each project below its minJobs; the hostShare cap; never-fits refusals per host and per fleet; the fleet view in doctor.
    3 Suggestions. Sizes per project from recorded peaks, shown in doctor, the panel and insights.
    4 Docs and release.
  2. 9 remaining items

  3. edgehero commented on Oct 7, 2026

    @edgehero
    OwnerAuthor

    Done in pi-dispatch 4.0.0 (https://github.com/edgehero/pi-dispatch/releases/tag/v4.0.0).

    What shipped, by phase

    Evidence

    • Measured on five setups (Docker Desktop, Docker 29, rootless Podman 4.9, Podman 5.8 rootful and rootless): each size read back exactly from the container's cgroup (memory, swap 0, CPU limit and weight, processes, shared memory); a job past its memory stopped alone while its neighbour ran on; two busy jobs split CPU in size order; with the CPU limit across jobs, busy jobs together stayed at the budget while a busy program outside kept its core.
    • Every phase went through review rounds (correctness, an adversarial reviewer, runs on the five setups), with the review scenarios kept as scripts and re-run after each fix. Every new test pin was checked by deliberately breaking the code.
    • The full suite passed on every merge (7272 tests at release), and all six required checks were green.

    Known limits, stated in the docs

    • Disk I/O, disk space and network bandwidth are not isolated between jobs.
    • On Linux with Docker's systemd driver or rootful Podman, the CPU limit across jobs needs one sudo command, which doctor prints; on rootful Podman and remote daemons doctor cannot read it back.
    • Suggestions read this machine's run records only, unless PI_LOGS_DIR is shared.
    • The panel's limit dialogs keep a project's size but do not prompt for one; a suggested size is applied by asking pi to run the call (it asks you to confirm).

    The machine-level view of how busy each host is, and was, is #599.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions