Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 13 additions & 0 deletions specs/constitution.md
Original file line number Diff line number Diff line change
Expand Up @@ -740,6 +740,18 @@ passing, on the record — issue #80.)
`repeat:`) and `removeJobScheduler` — or alert — past a threshold. **BullMQ will never do this for
us.** A cron silently re-running a wedging job is exactly the runaway `CONST-BUDGET-BEFORE-TOKENS`
exists to prevent, except unattended and overnight.
**A scheduled job that stalls AFTER its run was recorded is returned, never run.** A stall is not only a
dead worker. When Valkey is unreachable for longer than the lock renewal window while a run finishes,
the processor writes the run record, BullMQ refuses the completion ("Missing lock"), and the stall check
moves the scheduled job back to `wait`. Re-processing it is a second paid run of a job that already
ended, and its record overwrites the first. So the processor's first gate asks, for a job with
`stalledCounter > 0`, whether its record settles this attempt (`attempt` equal to `attemptsMade + 1`, an
outcome other than `failed`, started no earlier than the job less a 5 minute clock tolerance). If it
does, the job returns that recorded result without a container, a reservation or a new record, and
logs `job_lost_lock_after_completion`. No record (a worker that really died mid-run) keeps today's
behaviour, a record found and refused is logged as `job_lost_lock_record_rejected` before the job runs
again, and the per-scheduler stall guard still counts the stall either way, because the queue really
did lose the lock.
**The rule binds the HOST side too, and one whole class of refusal was on the wrong side of it until
issue #310.** A configuration fault is determinate by definition: the operator has to fix it, and the
same job with the same inputs refuses identically for as long as they do not. Such errors are tagged
Expand Down Expand Up @@ -1002,6 +1014,7 @@ passing, on the record — issue #80.)

| Date | Change |
|---|---|
| 2026-10-04 | Found this round (no issue): a scheduled job whose run finished while Valkey was unreachable longer than the lock renewal window was run again, paid. **`CONST-RETRY-INFRA-ONLY` AMENDED**, the Why only: a scheduled job that stalls after its completion was recorded is returned as its record says, without running (the processor's lost-lock gate: `stalledCounter > 0`, `attempt` equal to `attemptsMade + 1`, outcome not `failed`, started no earlier than the job less a 5 minute clock tolerance; a record found and refused is logged as `job_lost_lock_record_rejected`); a job with no such record runs as before, and the stall guard still counts the stall. The Statement and Acceptance are UNCHANGED, checked: a re-run of a completed job was a retry of a determinate outcome, which the Statement already forbids. **Code evidence**: worker/src/index.mjs -> makeProcessor (the first gate); worker/src/run-history.mjs -> makeSettledRecord. |
| 2026-10-04 | Issue #504, part A (delegated allocation: the doctrine and the pure modules). **No article changed.** **`CONST-BUDGET-BEFORE-TOKENS` UNCHANGED, checked**: the ordering is untouched; a delegated allocation changes only the values the dollar reserve compares against (a project's cap becomes the smaller of its row and its allocation), never when the reserve runs. **`CONST-ISOLATION-CONTAINER-PER-JOB` UNCHANGED, checked**: the operator's session still processes no adversarial input by premise; a model there that reads issue text breaks that premise, and `DES-DELEGATED-ALLOCATION-INSIDE-ENVELOPE` names it as a residual of the session tool, whose judgement belongs in a portfolio job (issue #505). |
| 2026-10-02 | Issue #501, parts 3 and 4, and #503 part 7 (dollar windows, reserved before the run and settled after it). **`CONST-BUDGET-BEFORE-TOKENS` AMENDED**: the Statement is the issue's: every spend cap is checked and reserved before a run, there are two ledgers (a job count and, when configured, dollars), a dollar reservation is the job's per-job cap held against every window that applies and replaced after the run by the metered cost or kept whole when the cost is not fully known, and the per-job cap is enforced in the container before each call; the constraint governs the ordering, the same for both ledgers. The Why gains why dollars reserve after both job-count reserves and before the container, why the amount is the per-job cap, why a refused dollar reservation is given back (the one departure from "a refused slot still counts"), and why a zero-rated local job reserves nothing yet keeps the ordering (a cap of 0 before each call). The Acceptance gains: given a dollar window is exhausted, a new trigger consumes zero provider tokens. **`CONST-RETRY-INFRA-ONLY` UNCHANGED, checked**: `dollar-cap` is a returned policy refusal; a Valkey fault in the dollar reserve is a never-started retry that refunds both ledgers. **Code evidence**: worker/src/dollar-budget.mjs; worker/src/processor.mjs -> runJob, isNeverStartedExit. The Statement says at least the reservation when the cost is not fully known, and that the per-job cap bounds what a job can spend before it runs, with the residuals named in `DES-DOLLAR-RESERVE-AND-SETTLE`, rather than that a job cannot spend more than it reserved. |
| 2026-10-02 | Issues #501 and #502, the shared seams for policy stops. **No article changed.** **`CONST-BUDGET-BEFORE-TOKENS` UNCHANGED, checked**: the new refusal of a cost cap or a model list the runner cannot enforce before a call is free and runs inside the container after the meter installs and before the session exists, so nothing is sent to a provider, and no worker gate moved. **`CONST-RETRY-INFRA-ONLY` UNCHANGED, checked**: that refusal is a determinate exit `2`, not retried. |
Expand Down
39 changes: 33 additions & 6 deletions specs/design.md
Original file line number Diff line number Diff line change
Expand Up @@ -6391,12 +6391,38 @@ a tunnel.
comment, bounded at one; a whole-deployment death between the terminal transition and the listener
loses both channels (a single crashed worker does NOT -- the stall path fails its job at next pickup
and the observing worker's listener fires); a graceful drain's exit can orphan an in-flight POST or
hook child. The DUPLICATION direction exists too, narrow and named (review finding): a worker that
comments a policy stop and then loses its job LOCK before BullMQ commits the return (crash, or a
partition outliving lock renewal) has its job stall-swept back to wait, and the re-pickup's
UnrecoverableError lands a second, `finishedOn`-guarded FAILED comment on the same issue -- a
pre-existing crash-consistency property of the queue that these comments make externally visible,
bounded at one duplicate, and not worth the idempotence store this entry already rejects. `OQ-023`'s
hook child. The DUPLICATION direction is now closed by the run record, which is the store that already
exists. A job that finishes and then loses its LOCK before BullMQ commits the return (a Valkey outage
longer than the lock renewal window, a partition, or a crash between the record and the commit) is taken
back by the stall check. Its completion is refused ("Missing lock"), and the next pickup fails it with
BullMQ's stall reason. That put the FAILED comment and the infra page on a job that had completed (or
on the same issue after a policy stop's own comment), and it ran a scheduled job again instead
(`CONST-RETRY-INFRA-ONLY`). So the failed listener, on a terminal failure whose message is exactly that
reason (`STALLED_FAILED_REASON` in `index.mjs`, pinned by a test against the installed bullmq's
`moveStalledJobsToWait`), first reads the job's record: this host's file, then the run mirror where one
is armed (`PI_WORKER_NAME`). When the record settles this attempt (same job id, `attempt` equal to
`attemptsMade`, which BullMQ has already raised by one, an outcome other than `failed`, and a `startedAt`
no earlier than the job's creation less `RECORD_CLOCK_SKEW_MS`, 5 minutes, so a reused id cannot match
an old record while a producer clock ahead of the worker's does not turn the check off), the terminal
line is `job_lost_lock_after_completion`, nothing is posted, and the hook fires only when the record is
a paid policy stop the completed listener would have paged for (that listener never saw the first
finish). A missing record, a mismatch or a lookup fault keeps the comment and the page, and a record
that was found and refused is logged as `job_lost_lock_record_rejected` with a fixed reason
(`unreadable`, `id-mismatch`, `other-attempt`, `failed`, `older-than-job`) and its source (`local` or
`mirror`; a mirror value that is there but is not a record is `unreadable` too), so an operator can see
why the comment was posted. The processor's gate says it ONCE per job id, stall count, attempt and
source (a bounded set of `REJECTED_SEEN_MAX`, 1000, oldest out first): it runs before every deferral
and BullMQ never resets a stall count, so a held job meets it on each pickup with the same verdict.
Residuals, each bounded at one comment
and one page unless said otherwise: on a fleet WITHOUT a shared `PI_LOGS_DIR`, a runtime-queue job
failed by a host that did not run it while the mirror holds no copy. The mirror write is not lost to
the outage itself (the shared client's offline queue sends it on reconnect, measured), so this is the
writing host exiting before Valkey returns (its offline queue is lost), or another host failing the job
before the writer has reconnected. A clock skew between producer and worker larger than the tolerance
(logged as `older-than-job`). A run that dies after its terminal comment but before its record. And a
suppressed job is NOT completed in the queue: it stays in BullMQ's failed set with the stall reason, so
the admin's views and `getJobCounts` show it failed while the worker log and the run record say it
completed; the record is the authority (`INT-RUN-HISTORY-FILE-CONTRACT`). `OQ-023`'s
prepare-policy silence stands untouched -- #288 is post-spend only, and that row's own hazard (a
`.pi/` cap breach commenting on every delivery) is exactly why.
- **Traces to**: `REQ-JOB-STATUS-COMMENTS`, `REQ-OPERATOR-FAILURE-NOTIFICATION`,
Expand Down Expand Up @@ -8169,3 +8195,4 @@ a tunnel.
| 2026-10-04 | Issue #504, part B. **`DES-DELEGATED-ALLOCATION-INSIDE-ENVELOPE` AMENDED**: part B is named as shipped (`worker/src/allocation.mjs`); the compare-and-set is spelled out (`CAS_SCRIPT` decodes with `cjson`, a null id is the empty string, the neutral seed is `SEED_SCRIPT`, which on a fresh seed sets `alloc:envelope:expected` to the seed's digest unconditionally in the same step, so a stale one left by a turn-off never survives a new seed, and on replacing a state that does not decode sets it only when absent, so a stale host cannot take the fleet through an unreadable plan); the audit order is split, the system's own changes (seed, expiry, re-base) written by the CAS winner after it won, a writer's change (plan, revert) written before the CAS with an `apply-failed` row after a lost one; the ladder gains `envelope-mismatch` after the writer rung and names `plan-invalid` before it; the host role in a re-base: every host reconciles at boot, on every envelope reload and at each pickup, so setting `alloc:envelope:expected` heals a mismatched host at its next job; `expected` names the APPLIED split's envelope (set by the seed's winner, by a host that agrees, and by a differing host that finds it absent), PR #574's review having found that the lazy seed let the first differing host re-base the fleet; the first host to reconcile at a first start or after the keys were deleted defines the split, and doctor names every host that differs; a changed-externally row is remembered only after it was written; a neutral state re-bases to the new envelope's neutral split, a plan keeps its id, `lastPlanAt` and expiry; a host with no envelope in a governed fleet refuses its jobs (one `EXISTS` per pickup), so delegation is turned off for a fleet by removing the envelope everywhere and deleting `alloc:plan` and `alloc:envelope:expected`; a fault while the pickup reads the split is `InfraRetry`, never a dropped job; two audit residuals are named (a system change whose row the file refuses stands, logged as `allocation_audit_row_lost` and still pushed to `alloc:log`; a writer's CAS that throws leaves its applied row with no `apply-failed` row); a stored state whose amounts are not safe non-negative integers is corrupt, one a newer build wrote is neither trusted nor overwritten; enforcement names the synthetic ledger and `_other`'s own `otherDollars` slot (a deviation from the plan's "through the project slot": a project the envelope does not name keeps its own row and counts in `_other`), with `capSource`; the containment check is wired with every load and reload; a local job's folder is judged where it is mounted: named inside a job path by text or by an ancestor's identity, a symlink farm refused with the configuration to use instead, the envelope's folder refused, a resolved folder in another project refused, a chained child enqueued on the named folder. Three Rejected entries: `_other` through the project slot, refusing a triggers reload whose new job path holds the envelope, and holding a local folder to the run roots by a flag set at enqueue. **`DES-DOLLAR-RESERVE-AND-SETTLE` AMENDED**: item 9 names `_other`'s ledger and `allocation-cap`; a new item 10, the worker half under an envelope: `min` per window, never replace; the synthetic project, repo and `_other` ledgers join the job-cost prefixes and settle through `holdPart`; `capSource` per ledger and `dollarCapSource` for the deployment decide `allocation-cap` against `dollar-cap` (a tie is the operator's); a shrink refuses new starts only; `envelope-mismatch` is a free gate before the mint. **`DES-FLEET-LEASES-FOR-SHARED-BOUNDS` AMENDED**: `alloc:lock` borrows the `SET NX PX` idiom and `RELEASE_IF_MINE` and bounds nothing; the leases are UNCHANGED, checked. **`DES-TERMINAL-COMMENTS-AND-FAILURE-HOOK`** UNCHANGED, checked: the new refusals comment once, pre-spend, and stay out of the hook's policy set. **Code evidence**: `worker/src/allocation.mjs`, `worker/src/processor.mjs` (the envelope gate, the resolved folder's project, `capSourceOf`), `worker/src/index.mjs` (`governedDollars` and the `InfraRetry` at pickup), `worker/src/start.mjs` (`reloadEnvelope`, `afterCommit`, `fpEnvelope`, the reaper), `worker/src/prepare-local.mjs` (`judgePlacement`), `worker/src/outbox.mjs`; tests `allocation-apply`, `allocation.integration`, `index-allocation`, `processor-allocation`, `prepare-local`, `outbox`, `start-wiring`, `doctor`. |
| 2026-10-04 | Issue #505, part A. **`DES-DELEGATED-ALLOCATION-INSIDE-ENVELOPE` AMENDED**: a portfolio job (`run.portfolio: true`, confirmed against the live triggers file at pickup) on a host whose envelope could not apply its plan (none, delegation off, or `portfolio-job` not a writer) is refused as `portfolio-no-envelope`, a free gate after `envelope-mismatch` and before the mint and the clone. Rejected: running it and recording the plan's refusal, which pays for a container whose one output is known to be refused. A job whose live entry lost the flag runs as an ordinary cron job. `DES-JOB-OUTBOX-CHAINING` and `DES-CRON-VIA-BULLMQ-SCHEDULER` UNCHANGED, checked (the snapshot and the plan channel are part B; a trigger fired by hand is one ordinary job under a `manual:` id). |
| 2026-10-04 | Issue #505, part B. **`DES-JOB-OUTBOX-CHAINING` AMENDED**: a second kind of request, the priorities plan in `/outbox/priorities.json`, collected on the completed-only path after the chain requests and enqueueing nothing, with the facts coming in as `/job/portfolio.json` and its own baked persona `guardrails/PORTFOLIO_PROTOCOL.md`. Rejected: a `type` discriminator in `request-<n>.json` (it would share the chain count cap, and an older worker would refuse it as `chain-bad-flow-name`), and Valkey, the envelope or an admin tool in the container. **`DES-DELEGATED-ALLOCATION-INSIDE-ENVELOPE` AMENDED**: the portfolio-job writer path (the snapshot, the collector, `recordRefusal` rows for the collector's own refusals, `plan-duplicate` on re-collection and `plan-too-soon` for a retried attempt, a race won by the first compare-and-set, a hand fire queued beside a tick and deferred behind it by the folder mutex, the rule that keeps a never-confirmed job's refusal out of the audit file and `alloc:log`, and `plan-collect-error` as an unknown outcome), and three rejected alternatives (the flag on job data alone, the live file alone at collection, an allocation module in the image). `REQ-AI-TRIGGERED-RUNS` UNCHANGED, checked: a plan enqueues nothing. |
| 2026-10-04 | Found this round (no issue): a job that completed while Valkey was unreachable was reported failed, or run again. **`DES-TERMINAL-COMMENTS-AND-FAILURE-HOOK` AMENDED**, Residuals: the duplication residual ("bounded at one duplicate, and not worth the idempotence store") is replaced by the record-read suppression. The run record is the existing store: on a terminal failure whose message is exactly BullMQ's stall reason (`STALLED_FAILED_REASON`, pinned against the installed bullmq), the failed listener reads the record (local file, then the run mirror on a declared fleet) and, when it settles this attempt (id, `attempt` equal to `attemptsMade`, outcome not `failed`, started no earlier than the job), logs `job_lost_lock_after_completion`, posts nothing, and pages only for a paid policy stop the completed listener would have paged for. `startedAt` may sit up to `RECORD_CLOCK_SKEW_MS` (5 minutes) before the job's creation, because the two instants come from two hosts' clocks; a record found and refused (a mirror value that is not a record included) is logged as `job_lost_lock_record_rejected` with a fixed reason and its source, by the processor's gate once per job id, stall count, attempt and source (a bounded set of 1000), since the gate runs before every deferral. The remaining residuals are named: a fleet without a shared logs dir whose mirror holds no copy (the writing host exits before Valkey returns, or another host fails the job before the writer reconnects; the offline queue otherwise delivers the mirror write), a skew beyond the tolerance, a run that dies between its comment and its record, and a suppressed job still sitting in BullMQ's failed set with the stall reason while the record says it completed. The Decision, Rejected list and Acceptance are UNCHANGED, checked. **Code evidence**: worker/src/start.mjs (the failed listener), worker/src/run-history.mjs (`makeReadRecord`, `recordSettlesAttempt`, `makeSettledRecord`), worker/src/run-mirror.mjs (`readMirroredRecord`), worker/src/index.mjs (`STALLED_FAILED_REASON`, the processor's lost-lock gate). |
Loading
Loading