Skip to content

fix(worker): map a job id onto the runtime's name rule, so a cron job can start a container (#435) - #436

Merged
edgehero merged 1 commit into
mainfrom
fix/435-cron-container-name
Sep 25, 2026
Merged

edgehero merged 1 commit into
mainfrom
fix/435-cron-container-name

Conversation

@edgehero

@edgehero edgehero commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

Closes #435.

A cron job could never start a container. BullMQ's job scheduler mints its id as repeat:<schedulerId>:<millis>, and jobContainerName used that id unchanged, on the strength of a comment saying BullMQ ids are already [A-Za-z0-9._-]. They are not for scheduler jobs, and the runtime refuses : in a container or network name.

Measured on rootless Podman 5.8.1:

  • pi-job-repeat:t:1790361600000 and its -net network are both refused at create (exit 125, "names must match [a-zA-Z0-9][a-zA-Z0-9_.-]*").
  • The sanitised name starts.
  • In a real worker, every cron attempt ended container-never-started before any spend.

Docker applies the same rule; that was not run here.

The fix: jobContainerName replaces every character outside [A-Za-z0-9._-] with _, as sanitizeJobId already does for file names. Every path that must find the container again goes through this function, so they all agree: the timeout's stop, the cancel, the per-job network and the log sink. Any id that is not a scheduler's is already of that shape and is unchanged; the test pins that beside the cron case, and checks the container and network names against the runtime's own rule. The mapping is not injective: repeat:a:1 and a literal repeat_a_1 share a name. The comment states this as a residual that nothing here reaches.

Specs: INT-CONTAINER-RUNTIME-CONTRACT amended, with a revision row. INT-EGRESS-POLICY-CONTRACT unchanged, checked.

Verified on a real host (Fedora 44, rootless Podman 5.8.1, the podman venue, a cron trigger firing every minute):

  • With the fix, each cron job started pi-job-repeat_<scheduler>_<millis>, with its -net network when egress was on and --network=private with PI_EGRESS=0. It reached the provider, got a 401 for a fake key, and its network and container were removed afterwards.
  • pi-dispatch cancel on a running cron job stopped the container by its sanitised name and recorded operator-cancel.
  • On main without the fix, the same jobs ended container-never-started (exit 125), with egress on and off.
  • Nothing compares a container name back to a job id. The cidfile and run-record paths never held the id's :, and the sandbox namespace is untouched.

CI-posture suite, rebased on main: 4393 tests, 0 fail, 1 skipped (systemd, on macOS). No version bump.

… can start a container (#435)

BullMQ's job scheduler mints a cron job's id as repeat:<schedulerId>:<millis>,
and jobContainerName used it unchanged, on the strength of a comment saying
BullMQ ids are already [A-Za-z0-9._-]. They are not for scheduler jobs, and
the runtime refuses ':' in a container or network name. Measured on rootless
Podman 5.8.1: pi-job-repeat:t:1790361600000 and its -net network are both
refused at create (exit 125, "names must match [a-zA-Z0-9][a-zA-Z0-9_.-]*"),
and in a real worker every cron attempt ended container-never-started before
any spend. Docker's rule is the same; not run here.

jobContainerName now replaces every character outside [A-Za-z0-9._-] with
'_', as sanitizeJobId does for file names. Every path that must find the
container again (the timeout's stop, the cancel, the per-job network, the
log sink) asks this function, so they all agree. Every id other than a
scheduler's is already of that shape and is unchanged, which the test pins
beside the cron case; the result is checked against the runtime's own rule
for the container and its network. Not injective (repeat:a:1 and a literal
repeat_a_1 share a name), stated in the comment as a residual no producer
here reaches.

INT-CONTAINER-RUNTIME-CONTRACT amended, with a revision row.
INT-EGRESS-POLICY-CONTRACT UNCHANGED, checked.

Signed-off-by: Rob Boerman <robboerman@live.nl>
@edgehero
edgehero force-pushed the fix/435-cron-container-name branch from 056dbd8 to 477b54b Compare September 25, 2026 19:10
@edgehero
edgehero marked this pull request as ready for review September 25, 2026 19:10
@edgehero
edgehero merged commit 8db8cd7 into main Sep 25, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Cron jobs never start: the scheduler's repeat:<id>:<millis> job id is not a valid container or network name

1 participant