fix(worker): map a job id onto the runtime's name rule, so a cron job can start a container (#435) - #436
Merged
Conversation
… can start a container (#435) BullMQ's job scheduler mints a cron job's id as repeat:<schedulerId>:<millis>, and jobContainerName used it unchanged, on the strength of a comment saying BullMQ ids are already [A-Za-z0-9._-]. They are not for scheduler jobs, and the runtime refuses ':' in a container or network name. Measured on rootless Podman 5.8.1: pi-job-repeat:t:1790361600000 and its -net network are both refused at create (exit 125, "names must match [a-zA-Z0-9][a-zA-Z0-9_.-]*"), and in a real worker every cron attempt ended container-never-started before any spend. Docker's rule is the same; not run here. jobContainerName now replaces every character outside [A-Za-z0-9._-] with '_', as sanitizeJobId does for file names. Every path that must find the container again (the timeout's stop, the cancel, the per-job network, the log sink) asks this function, so they all agree. Every id other than a scheduler's is already of that shape and is unchanged, which the test pins beside the cron case; the result is checked against the runtime's own rule for the container and its network. Not injective (repeat:a:1 and a literal repeat_a_1 share a name), stated in the comment as a residual no producer here reaches. INT-CONTAINER-RUNTIME-CONTRACT amended, with a revision row. INT-EGRESS-POLICY-CONTRACT UNCHANGED, checked. Signed-off-by: Rob Boerman <robboerman@live.nl>
edgehero
force-pushed
the
fix/435-cron-container-name
branch
from
September 25, 2026 19:10
056dbd8 to
477b54b
Compare
edgehero
marked this pull request as ready for review
September 25, 2026 19:10
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #435.
A cron job could never start a container. BullMQ's job scheduler mints its id as
repeat:<schedulerId>:<millis>, andjobContainerNameused that id unchanged, on the strength of a comment saying BullMQ ids are already[A-Za-z0-9._-]. They are not for scheduler jobs, and the runtime refuses:in a container or network name.Measured on rootless Podman 5.8.1:
pi-job-repeat:t:1790361600000and its-netnetwork are both refused at create (exit 125, "names must match[a-zA-Z0-9][a-zA-Z0-9_.-]*").container-never-startedbefore any spend.Docker applies the same rule; that was not run here.
The fix:
jobContainerNamereplaces every character outside[A-Za-z0-9._-]with_, assanitizeJobIdalready does for file names. Every path that must find the container again goes through this function, so they all agree: the timeout's stop, the cancel, the per-job network and the log sink. Any id that is not a scheduler's is already of that shape and is unchanged; the test pins that beside the cron case, and checks the container and network names against the runtime's own rule. The mapping is not injective:repeat:a:1and a literalrepeat_a_1share a name. The comment states this as a residual that nothing here reaches.Specs:
INT-CONTAINER-RUNTIME-CONTRACTamended, with a revision row.INT-EGRESS-POLICY-CONTRACTunchanged, checked.Verified on a real host (Fedora 44, rootless Podman 5.8.1, the
podmanvenue, a cron trigger firing every minute):pi-job-repeat_<scheduler>_<millis>, with its-netnetwork when egress was on and--network=privatewithPI_EGRESS=0. It reached the provider, got a 401 for a fake key, and its network and container were removed afterwards.pi-dispatch cancelon a running cron job stopped the container by its sanitised name and recordedoperator-cancel.mainwithout the fix, the same jobs endedcontainer-never-started(exit 125), with egress on and off.:, and the sandbox namespace is untouched.CI-posture suite, rebased on main: 4393 tests, 0 fail, 1 skipped (systemd, on macOS). No version bump.