Describe the bug
On a customer-managed fleet with a Linux worker, a fleet hostConfiguration script can never run. The worker agent service runs as the unprivileged deadline-worker-agent user, but HostConfigurationScriptRunner hardcodes runas_user=PosixSessionUser(user="root"), so the script is invoked through sudo -u root. install-deadline-worker does not create any sudoers rule permitting that escalation, so sudo refuses.
The result is not a degraded script run, it is a crash loop: host configuration exits non-zero, the agent logs Cannot run jobs, exiting and exits, the service manager restarts it, and the cycle repeats every 2 to 3 seconds indefinitely. The worker never reaches a usable state and never picks up a job.
This appears to affect every customer-managed Linux fleet that uses a host configuration script with a default agent install.
Relevant code:
src/deadline_worker_agent/startup/host_configuration_script.py, the runas_user default on HostConfigurationScriptRunner.__init__:
runas_user=PosixSessionUser(user="root") if sys.platform != "win32" else None,
src/deadline_worker_agent/installer/, which creates the agent user and the systemd unit but adds no sudoers entry for root escalation. The only sudoers rule the installer manages is the optional shutdown rule behind --allow-shutdown.
Expected Behaviour
After install-deadline-worker ... --start on a Linux host, a fleet host configuration script should execute with the elevated privileges it is documented to get, and the worker should proceed to STARTED and accept jobs.
Whichever way it is resolved, one of these should hold:
- The installer provisions the sudoers rule that host configuration requires, the same way it already does for
--allow-shutdown; or
- The requirement is documented, with the exact rule an operator must add, and the agent fails with an actionable message naming the missing sudoers entry rather than surfacing raw
sudo output; or
- The agent detects at startup that it cannot escalate and refuses to enter the retry loop, reporting the cause once.
Current Behaviour
Every host configuration attempt fails with sudo's refusal, and the agent crash-loops. Abridged from /var/log/amazon/deadline/worker-agent.log:
[INFO ] --------- Running Host Configuration Script ---------
[INFO ] Running command /usr/bin/sudo -u root -i /usr/bin/setsid -w /var/lib/deadline/tmpikj83vsl.sh
[WARNING ] Unable to determine signal target: unable to detect subprocess before timeout
[INFO ] Command started as pid: 7692
[INFO ] Output:
[INFO ] sudo: I'm sorry deadline-worker-agent. I'm afraid I can't do that
[INFO ] Process pid 7692 exited with code: 1 (unsigned) / 0x1 (hex)
[INFO ] --------- Finished running Host Configuration Script, exit code: 1 ---------
[CRITICAL] Worker.HostConfiguration Worker Agent host configuration failed with exit code 1. Cannot run jobs, exiting.
That block repeats about every 2 to 3 seconds. In one minute of uptime the worker's CloudWatch log stream contained 30 complete host configuration cycles (60 Worker/HostConfiguration events, 30 INFO and 30 CRITICAL), plus 630 untyped output events.
Three things make this harder to diagnose than it needs to be:
/etc/sudoers.d/ contains no agent entry after a default install, so there is no hint that a sudoers rule was ever expected.
- The retry has no backoff, so the underlying one-line cause is buried under hundreds of repetitions.
- The
sudo refusal is logged at INFO as untyped output, so filtering the worker log by level or by subtype: HostConfiguration does not surface the actual reason.
Reproduction Steps
-
Create a customer-managed fleet with a Linux workerCapabilities (osFamily: linux, cpuArchitectureType: x86_64) and attach a host configuration script. Any script will do, since it never executes:
{"scriptBody": "#!/usr/bin/env bash\necho hello from host config\n", "scriptTimeoutSeconds": 300}
-
On a fresh Ubuntu host, install and start the agent with default flags:
python3 -m venv /opt/deadline-venv
/opt/deadline-venv/bin/pip install deadline-cloud-worker-agent
/opt/deadline-venv/bin/install-deadline-worker \
--farm-id <farm-id> --fleet-id <fleet-id> \
--region <region> --start --yes
-
Wait about 30 seconds, then inspect the log:
grep -c "Running Host Configuration Script" /var/log/amazon/deadline/worker-agent.log
grep "sudo: I'm sorry" /var/log/amazon/deadline/worker-agent.log | head
ls -la /etc/sudoers.d/
Expected: the script runs once and the worker starts. Actual: dozens of attempts, each failing on the sudo refusal, and no agent entry in /etc/sudoers.d/.
-
Confirm the cause by granting the escalation the runner needs (permitting the agent user to run as root) and restarting deadline-worker. The script then executes and the worker starts normally.
Note that this cannot be reproduced on a service-managed fleet, since the agent there ships in the service-managed AMI with its privileges already arranged.
Environment
- Operating system: Ubuntu 26.04 LTS, amd64, on EC2
- Version of this package:
deadline-cloud-worker-agent 0.33.3
Other details:
- Python 3.14 (distro
python3 in a venv at /opt/deadline-venv)
openjd.sessions 0.12.0, openjd.model 0.11.14, deadline.job_attachments 0.1.3
- Fleet type: customer-managed,
mode: NO_SCALING, minWorkerCount 0, maxWorkerCount 1
- Installed with default flags, so the agent user is
deadline-worker-agent and allow_ec2_instance_profile is True
- Service manager: systemd, unit
deadline-worker.service as created by the installer
/etc/sudoers.d/ contained only the cloud-init and ssm-agent entries before and after installation
Describe the bug
On a customer-managed fleet with a Linux worker, a fleet
hostConfigurationscript can never run. The worker agent service runs as the unprivilegeddeadline-worker-agentuser, butHostConfigurationScriptRunnerhardcodesrunas_user=PosixSessionUser(user="root"), so the script is invoked throughsudo -u root.install-deadline-workerdoes not create any sudoers rule permitting that escalation, sosudorefuses.The result is not a degraded script run, it is a crash loop: host configuration exits non-zero, the agent logs
Cannot run jobs, exitingand exits, the service manager restarts it, and the cycle repeats every 2 to 3 seconds indefinitely. The worker never reaches a usable state and never picks up a job.This appears to affect every customer-managed Linux fleet that uses a host configuration script with a default agent install.
Relevant code:
src/deadline_worker_agent/startup/host_configuration_script.py, therunas_userdefault onHostConfigurationScriptRunner.__init__:src/deadline_worker_agent/installer/, which creates the agent user and the systemd unit but adds no sudoers entry for root escalation. The only sudoers rule the installer manages is the optional shutdown rule behind--allow-shutdown.Expected Behaviour
After
install-deadline-worker ... --starton a Linux host, a fleet host configuration script should execute with the elevated privileges it is documented to get, and the worker should proceed toSTARTEDand accept jobs.Whichever way it is resolved, one of these should hold:
--allow-shutdown; orsudooutput; orCurrent Behaviour
Every host configuration attempt fails with
sudo's refusal, and the agent crash-loops. Abridged from/var/log/amazon/deadline/worker-agent.log:That block repeats about every 2 to 3 seconds. In one minute of uptime the worker's CloudWatch log stream contained 30 complete host configuration cycles (60
Worker/HostConfigurationevents, 30INFOand 30CRITICAL), plus 630 untyped output events.Three things make this harder to diagnose than it needs to be:
/etc/sudoers.d/contains no agent entry after a default install, so there is no hint that a sudoers rule was ever expected.sudorefusal is logged atINFOas untyped output, so filtering the worker log by level or bysubtype: HostConfigurationdoes not surface the actual reason.Reproduction Steps
Create a customer-managed fleet with a Linux
workerCapabilities(osFamily: linux,cpuArchitectureType: x86_64) and attach a host configuration script. Any script will do, since it never executes:{"scriptBody": "#!/usr/bin/env bash\necho hello from host config\n", "scriptTimeoutSeconds": 300}On a fresh Ubuntu host, install and start the agent with default flags:
Wait about 30 seconds, then inspect the log:
Expected: the script runs once and the worker starts. Actual: dozens of attempts, each failing on the
sudorefusal, and no agent entry in/etc/sudoers.d/.Confirm the cause by granting the escalation the runner needs (permitting the agent user to run as root) and restarting
deadline-worker. The script then executes and the worker starts normally.Note that this cannot be reproduced on a service-managed fleet, since the agent there ships in the service-managed AMI with its privileges already arranged.
Environment
deadline-cloud-worker-agent0.33.3Other details:
python3in a venv at/opt/deadline-venv)openjd.sessions0.12.0,openjd.model0.11.14,deadline.job_attachments0.1.3mode: NO_SCALING,minWorkerCount0,maxWorkerCount1deadline-worker-agentandallow_ec2_instance_profileisTruedeadline-worker.serviceas created by the installer/etc/sudoers.d/contained only the cloud-init and ssm-agent entries before and after installation