Skip to content
Open
Show file tree
Hide file tree
Changes from 8 commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
137 changes: 137 additions & 0 deletions src/deadline/job_attachments/_system_commands.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,137 @@
# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved.

"""Resolution of system command names to absolute paths, without consulting PATH.

The problem: the VFS mount and unmount paths invoke ``sudo`` to act as the job
user, and query mount state with ``findmnt``. Invoking those by bare name resolves
them through ``PATH``, which makes the binary actually run depend on the search
path of whatever launched the process, for commands that cross a user boundary.

The solution: callers pass a bare name here and get back an absolute path found by
scanning a fixed list of trusted directories, so ``PATH`` plays no part.

Three properties make that work, and all three are easy to undo by accident:

* ``PATH`` is never read. Not directly, and not through :func:`shutil.which`,
which resolves via ``PATH`` and so would restore the original behaviour while
looking like a fix.
* Only paths under :data:`TRUSTED_SYSTEM_DIRECTORIES` are returned. A name
containing a path separator is rejected, because ``os.path.join`` would
otherwise let ``../../tmp/evil`` escape the directory being searched.
* A missing command raises. Returning the bare name as a fallback would put
resolution back on ``PATH`` while the code still read as though it did not.

A resolver rather than absolute-path literals, because the locations are not
universal: NixOS keeps the setuid ``sudo`` wrapper at ``/run/wrappers/bin/sudo``,
so a hardcoded ``/usr/bin/sudo`` would leave those hosts unable to mount at all.
"""

from __future__ import annotations

import os as _os
from typing import Optional as _Optional, Tuple as _Tuple

__all__ = [
"SystemCommandNotFoundError",
"TRUSTED_SYSTEM_DIRECTORIES",
"find_system_command",
"system_command_path",
]


TRUSTED_SYSTEM_DIRECTORIES: _Tuple[str, ...] = (
# Ordered, deliberately. On NixOS the setuid `sudo` wrapper lives here and the
# /usr/bin copy is absent or not setuid, so this must be searched first.
"/run/wrappers/bin",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two things about the trusted-directory list worth a second look, since this module's entire value rests on the claim that these locations are trustworthy:

  1. /run/wrappers/bin is searched first on every platform, not just NixOS. On a non-NixOS host this directory normally does not exist, so the entry is inert — but it is unconditionally given priority over /usr/bin. /run is a root-owned tmpfs on typical distributions, so exploiting it needs root already; still, giving highest precedence to a path that is non-standard on the vast majority of target hosts is the opposite of the "fixed list of trusted absolute directories" premise. Consider gating it (e.g. only prepend when /run/wrappers/bin/sudo is a setuid file, or when a NixOS marker such as /etc/NIXOS is present), or at minimum documenting that the entry is expected to be absent elsewhere.

  2. No ownership or write-permission check on the resolved directory/file. _is_executable_file only checks isfile + X_OK. A resolver whose stated purpose is defeating untrusted search paths would normally also refuse a candidate whose directory or file is group/world-writable, or not root-owned — otherwise a misconfigured /usr/local-style directory (or a symlink planted inside a searched directory) is trusted purely because of where it appears in the list. If that check is deliberately out of scope, saying so in the docstring alongside the three properties already listed would make the boundary explicit.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both points correct, and both parked rather than fixed. On the ordering: /run/wrappers/bin does get priority on every platform though it is NixOS-specific, and it is normally absent, costing one stat. On trust: _is_executable_file checks only isfile plus an execute bit, never root ownership or group and world writability, so membership is positional. Parked because reaching the exposure needs root or equivalent already, and an ownership check has the same failure shape as the X_OK check that already caused one regression in this series by testing permissions as the wrong user. Recorded so it is not lost, and the module docstring no longer implies a stronger guarantee than the code provides.

# ...and these two NixOS entries are a pair. /run/wrappers/bin holds only the
# setuid/setcap wrappers, so on NixOS it resolves `sudo` and nothing else:
# /usr/bin holds just `env`, /bin just `sh`, and the sbin directories are
# absent. `findmnt` lives in this symlink farm, which nixos-rebuild manages and
# root owns, so it is trust-equivalent to /usr/bin there. Without it the
# ordering above would resolve `sudo` and then fail on `findmnt`.
"/run/current-system/sw/bin",
"/usr/bin",
"/bin",
# sbin last: on non-usr-merged distributions some system commands exist only
# under /sbin.
"/usr/sbin",
"/sbin",
)


class SystemCommandNotFoundError(Exception):
"""A required system command was not present in any trusted directory.

Deliberately not a :class:`FileNotFoundError`, because ``vfs`` already uses that
type for something else. Three ``except FileNotFoundError`` blocks there mean
"the VFS pid file is missing", and one of them wraps a call chain that reaches
this resolver: ``kill_all_processes`` -> ``shutdown_libfuse_mount`` ->
``wait_for_mount`` -> ``is_mount``. Inheriting from ``FileNotFoundError`` let a
resolution failure be reported as a missing pid file and skip the cleanup that
follows.

Callers in ``vfs`` translate this into :class:`VFSExecutableMissingError`, which
is the type their own callers already handle by falling back to a copy-based
sync. This differs from the sibling resolver in ``openjd-sessions``, where
inheriting from ``OSError`` is correct because its cancel path deliberately
catches ``OSError`` so a failed signal cannot unwind a cancelation. Same
problem, opposite answer, because the surrounding handlers differ.
"""


def _validate_command_name(name: str) -> None:
"""Reject anything that is not a bare command name."""
if not name:
raise ValueError("A system command name must not be empty.")
if name in (_os.curdir, _os.pardir):
raise ValueError(f"{name!r} is not a system command name.")
# Both separators are checked on both platforms. A backslash is a legal POSIX
# filename character, but no command resolved here contains one, and treating
# it as suspect keeps the check identical rather than subtly weaker on POSIX.
# The colon is rejected for the same reason, and it is not hypothetical:
# ntpath.join(r"C:\Windows\System32", "D:evil") == "D:evil". A drive-relative
# name discards the trusted prefix while containing no separator at all, so a
# separator-only check lets it through. posixpath joins it harmlessly, but the
# guard belongs here rather than depending on which os.path is loaded.
if "/" in name or "\\" in name or ":" in name:
raise ValueError(
f"A system command name must not contain a path separator or drive "
f"specifier, but got {name!r}."
)


def _is_executable_file(path: str) -> bool:
return _os.path.isfile(path) and _os.access(path, _os.X_OK)


def find_system_command(name: str) -> _Optional[str]:
"""Return the absolute path to ``name``, or ``None`` if it is not installed.

``PATH`` is not consulted. Use this when the command's absence is tolerable;
use :func:`system_command_path` when it is required.

Raises:
ValueError: if ``name`` is not a bare command name.
"""
_validate_command_name(name)
for directory in TRUSTED_SYSTEM_DIRECTORIES:
candidate = _os.path.join(directory, name)
if _is_executable_file(candidate):
return candidate
return None


def system_command_path(name: str) -> str:
"""Return the absolute path to ``name``.

Raises:
ValueError: if ``name`` is not a bare command name.
SystemCommandNotFoundError: if ``name`` is in no trusted directory.
"""
path = find_system_command(name)
if path is None:
raise SystemCommandNotFoundError(
f"Could not find the system command {name!r} in any trusted directory "
f"({', '.join(TRUSTED_SYSTEM_DIRECTORIES)}). PATH is deliberately not searched."
)
return path
63 changes: 60 additions & 3 deletions src/deadline/job_attachments/vfs.py
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,15 @@
import threading
from typing import Callable, Dict, Union, Optional

# Aliased under private names. Imported plainly, these two become public
# attributes of `deadline.job_attachments.vfs`, which the API-surface check
# reports as added public aliases -- an unintended widening of the package's
# contract from what is meant to be an internal helper. All three are aliased,
# including the exception: re-exporting it plainly failed the same gate a second
# time, for the same reason.
from ._system_commands import SystemCommandNotFoundError as _SystemCommandNotFoundError
from ._system_commands import find_system_command as _find_system_command
from ._system_commands import system_command_path as _system_command_path
from .exceptions import (
VFSExecutableMissingError,
VFSFailedToMountError,
Expand Down Expand Up @@ -119,7 +128,24 @@
if not os.path.exists(fusermount3_path):
log.warning(f"fusermount3 not found at {cls.find_vfs_link_dir()}")
return None
return ["sudo", "-u", os_user, fusermount3_path, "-u", mount_path]
# find_system_command rather than system_command_path: this function's
# contract is Optional[list], and it already answers "a binary I need is
# missing" with a warning and None just above. Raising for sudo instead
# would be a second, undeclared failure mode on the same line of code --
# and it escapes into shutdown_libfuse_mount's cleanup path, whose only
# handling for this function is the None check.
sudo_path = _find_system_command("sudo")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The rationale recorded here — and the contract the new TestGetShutdownArgsFailureContract class pins — is that get_shutdown_args answers "a binary I need is missing" with None and never raises, so nothing escapes into shutdown_libfuse_mounts cleanup path. Good goal, but the function already has a raising path two lines above, on the very line the comment describes as the one that would gain "a second, undeclared failure mode":

fusermount3_path = os.path.join(cls.find_vfs_link_dir(), "fusermount3")   # vfs.py:127

find_vfs_link_dir() (vfs.py:326) calls find_vfs(), which raises VFSExecutableMissingError when the VFS executable cannot be located (vfs.py:412-465). So "the VFS binary is missing" already propagates out of get_shutdown_args as an exception, while "sudo is missing" and "fusermount3 is missing" return None. Three missing-binary cases, two different answers.

That matters beyond documentation accuracy, because one caller is mid-rewrite of the pid file when it happens. kill_process_at_mount (vfs.py:188) has already opened the file with "w" — truncating it — and is writing surviving entries back inside the loop when it calls shutdown_libfuse_mount at vfs.py:196. An exception raised from find_vfs_link_dir() at that point exits the with block with only the entries written so far, and the except FileNotFoundError at vfs.py:200 does not catch it, so it propagates through download.handle_existing_vfs (download.py:1172) and out of mount_vfs_from_manifests — where asset_sync catches VFSExecutableMissingError and quietly falls back to COPIED. Net result: a truncated pid file plus a silent fallback, with the mounts still up.

Two ways to make the stated contract true rather than half-true:

  • Wrap the find_vfs_link_dir() call in the same warn-and-return-None shape as the two checks below it, so all three missing-binary cases exit the function identically.
  • Or narrow the claim: TestGetShutdownArgsFailureContracts docstring says "get_shutdown_args must answer a missing binary with None, not an exception", which reads as a general property when the test only covers sudo. Saying "a missing sudo or fusermount3" and noting find_vfs_link_dir as a known raising path would keep the reasoning honest for whoever next picks a resolver here.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right on the facts, and the comment was overclaiming. Fixed by narrowing the claim rather than by changing behaviour: the comment and TestGetShutdownArgsFailureContract's docstring are now scoped to sudo and fusermount3, and both name find_vfs_link_dir() -> find_vfs() as a known raising path, so the asymmetry is stated instead of implicitly denied.

The second consequence no longer applies. kill_process_at_mount now reads the pid file fully and writes once at the end, as part of a separate fix in this revision for entries being dropped after a failed unmount, so an exception from find_vfs_link_dir() can no longer leave the file truncated mid-rewrite.

Making all three cases answer identically is a behaviour change to a pre-existing path; parked as item 9.

if sudo_path is None:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new return None here is reached in cases the old code never returned None for, and the consequence downstream is that pid-file state is discarded while the mounts are still up.

kill_all_processes (vfs.py:108-117) ignores the return value of shutdown_libfuse_mount and unconditionally runs os.remove(pid_file_path) afterwards. kill_process_at_mount (vfs.py:186-198) is worse: it sets mount_point_found = True and skips file.write(line) for the matched entry, so the entry is dropped from the pid file regardless of whether the unmount actually happened.

With bare "sudo", a host without sudo produced FileNotFoundError from subprocess.run inside shutdown_libfuse_mount, which is not caught there (only CalledProcessError is) and so propagated up to the except FileNotFoundError in kill_all_processes — reported as a missing pid file, but crucially before os.remove, so the pid file survived and the entries were still there for a later attempt.

Now the same host takes the new branch: warn, None, False, loop continues, pid file deleted. The mounts are still mounted, and the only record of which mount points and pids existed is gone, so nothing can clean them up afterwards. AssetSync.cleanup_session (asset_sync.py:1086-1091) sees no exception at all and reports success.

Two hosts this is reachable on today: any container/image without sudo installed, and any install that places sudo outside TRUSTED_SYSTEM_DIRECTORIES (e.g. /usr/local/bin/sudo, which is where MacPorts and some FreeBSD-derived layouts put it).

The non-raising choice for this function is well-argued in the comment above; the gap is that neither caller does anything with the False. Propagating the failure so kill_all_processes skips os.remove when any mount failed to shut down, and kill_process_at_mount writes the line back, would keep the return contract and stop losing the state.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct, and deferred with the other two. kill_all_processes ignoring the return value and removing the pid file regardless is the part that turns this into lost state rather than a failed cleanup, and kill_process_at_mount setting mount_point_found = True compounds it.

Both of those behaviours predate this change; what is new is a path that reaches them. Fixing it properly means changing how those two functions treat a failed shutdown, which is beyond command resolution and belongs with the owners of that cleanup logic.

log.warning("sudo not found in any trusted directory; cannot unmount as the job user")
return None
return [
sudo_path,
"-u",
os_user,
fusermount3_path,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

High-level note on the scope of this hardening: the two commands resolved through the new trusted resolver (sudo, findmnt) are the ones that were least influenced by less-trusted input, while the binaries actually executed under sudo are still resolved via PATH.

  • fusermount3_path here comes from find_vfs_link_dir() → find_vfs(), which starts with shutil.which(DEADLINE_VFS_EXECUTABLE) (vfs.py:357) and then falls back to $DEADLINE_VFS_INSTALL_PATH/cwd-relative bin/deadline_vfs. That path is then passed to sudo -u <os_user>.
  • build_launch_command (vfs.py:299) resolves the launch script from $DEADLINE_VFS_INSTALL_PATH and runs it via sudo -E -u <os_user>.

So if the threat model is "part of the search path is influenced by less-trusted input" (CWE-426, as the new module docstring states), pinning sudo while leaving the sudo argument resolved via PATH/env-var/cwd leaves the higher-impact half of the vector in place — -E even forwards the environment through.

Not asking to expand this PR necessarily, but it would be worth stating in the PR description / module docstring that find_vfs() and find_vfs_launch_script() are knowingly out of scope, so this is not later read as "the untrusted-search-path issue in vfs.py is closed."

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fair, and now stated explicitly in 4d82305 rather than left implicit. find_vfs carries a note that it deliberately does not use the resolver, with the reasoning: the resolver is for system commands whose locations are fixed and few, whereas the VFS executable is a shipped artifact whose location is a deployment choice, which is why shutil.which, $DEADLINE_VFS_INSTALL_PATH and a cwd-relative path are all consulted. Narrowing that is a question about how the VFS is deployed, not about command resolution, and it is knowingly out of scope here.

"-u",
mount_path,
]

@classmethod
def shutdown_libfuse_mount(cls, mount_path: str, os_user: str, session_dir: Path) -> bool:

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new comment above (vfs.py:131-136) rests on the premise that shutdown_libfuse_mount's cleanup path must not throw, and that its "only handling for this function is the None check." That premise is right, but this method already violates it independently of the resolver — and it is worth knowing before more reasoning is layered on top of it:

try:
    run_result = subprocess.run(shutdown_args, check=True)
except subprocess.CalledProcessError as e:
    log.warning(f"Shutdown failed with error {e}")
    # Don't reraise, check if mount is gone
log.info(f"Shutdown returns {run_result.returncode}")   # vfs.py:166

When check=True raises, run_result is never bound, so line 166 raises UnboundLocalError — precisely defeating the # Don't reraise, check if mount is gone intent one line above, and never reaching wait_for_mount. fusermount3 -u exits nonzero for ordinary reasons (target not mounted, device busy, permission denied for the job user), so this is not an exotic path.

Downstream that lands in the same place as the other cleanup concerns on this PR: kill_all_processes (vfs.py:109-118) only catches FileNotFoundError, so an UnboundLocalError aborts the loop over remaining mounts and skips os.remove(pid_file_path); AssetSync.cleanup_session (asset_sync.py:1086-1091) only catches VFSExecutableMissingError, so it propagates out of session cleanup entirely.

This is pre-existing rather than introduced here, so it is fair to leave out of scope. But the PR is specifically reasoning about which exceptions can escape this function, and run_result is the one that already does — so either fixing it (initialise run_result = None and branch, or move the log.info into an else:) or noting it as known-and-separate would keep that reasoning accurate.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, and thank you for checking the premise rather than the conclusion. subprocess.run(..., check=True) does raise CalledProcessError out of shutdown_libfuse_mount, and run_result can be unbound on that path, so the non-throwing contract my comment appealed to is not one this method keeps today.

That is pre-existing and untouched by this change, so I have not folded a fix into it. The comment it undercuts has been narrowed accordingly in the module note rather than left asserting more than the code delivers.

Expand Down Expand Up @@ -211,7 +237,24 @@
os.path.ismount returns false for libfuse mounts owned by "other users",
use findmnt instead
"""
return subprocess.run(["findmnt", path]).returncode == 0
return subprocess.run([cls._resolve_or_raise("findmnt"), path]).returncode == 0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is_mount raising here defeats the non-raising contract that get_shutdown_args was deliberately written to uphold, because the two sit on the same call chain.

get_shutdown_args (vfs.py:137) uses the non-raising _find_system_command, and the comment above it gives the reason: an exception "escapes into shutdown_libfuse_mount's cleanup path, whose only handling for this function is the None check." But shutdown_libfuse_mount ends with:

return cls.wait_for_mount(mount_path, session_dir, expected=False)   # vfs.py:164

and wait_for_mount -> is_mount -> _resolve_or_raise("findmnt") raises VFSExecutableMissingError. So shutdown_libfuse_mount can still throw out of the exact cleanup path the get_shutdown_args comment says must not throw — just via findmnt instead of sudo.

Concretely, in kill_all_processes (vfs.py:108-117) that exception is raised from inside the for line in file.readlines() loop, so it aborts before the remaining mounts are processed and skips os.remove(pid_file_path). It then surfaces in AssetSync.cleanup_session (asset_sync.py:1090) as Virtual File System not found, no processes to kill — which is inaccurate: processes existed, some may have been killed, and the pid file is still on disk.

This is reachable whenever sudo resolves but findmnt does not — e.g. a util-linux install that only provides /usr/local/sbin/findmnt, since /usr/local/* is (intentionally) not in TRUSTED_SYSTEM_DIRECTORIES. Previously the bare subprocess.run(["findmnt", path]) raised FileNotFoundError, which the surrounding except FileNotFoundError swallowed, so this path did not propagate out of kill_all_processes at all.

If the intent is that is_mount may raise, that is defensible — but then the reasoning recorded at vfs.py:129-135 is only half-true and the shutdown path is not actually exception-free. The alternative is to make is_mount mirror get_shutdown_args: resolve with _find_system_command and treat an unresolvable findmnt as "cannot determine mount state" (return False, with a warning) so unmount cleanup degrades the same way it does for a missing sudo or fusermount3.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct, and deferred. The two do sit on the same call chain, and the asymmetry is real: get_shutdown_args returns None while is_mount raises, so the contract I wrote a comment about is not one the module actually keeps end to end.

Deferring because it is not reachable on a supported host. findmnt is /usr/bin/findmnt everywhere this runs, so the raising path needs a host where it is absent from all six trusted directories. Making the two consistent means choosing which way, and every option changes behaviour on a cleanup path beyond command resolution: either is_mount starts returning a bool that conflates "not mounted" with "cannot tell", or the mount and unmount paths get different failure contracts on purpose. That is a design call for this module's owners rather than something to settle at the end of a security review.


@staticmethod
def _resolve_or_raise(name: str) -> str:
"""Resolve a system command, reporting failure as VFSExecutableMissingError.

The translation is the point. Callers of the mount and unmount paths handle
``VFSExecutableMissingError`` -- ``asset_sync`` catches it and falls back to a
copy-based sync -- and handle no other type for "a binary I need is not
here". Letting the resolver's own exception escape would either bypass that
fallback, or, if it inherited from ``FileNotFoundError``, be misreported by
the ``except FileNotFoundError`` blocks in this module that mean "the VFS pid
file is missing".
"""
try:
return _system_command_path(name)
except _SystemCommandNotFoundError as e:
raise VFSExecutableMissingError(str(e)) from e

@classmethod
def wait_for_mount(cls, mount_path, session_dir, mount_wait_seconds=60, expected=True) -> bool:
Expand Down Expand Up @@ -268,7 +311,7 @@
log_file_path = self.logs_folder_path(session_dir) / log_file_name
log.log(log_level, f"Printing last {lines} lines from {log_file_path}")
if not os.path.exists(log_file_path):
log.warning(f"No log file found at {log_file_path}")

Check failure

Code scanning / CodeQL

Clear-text logging of sensitive information High

This expression logs
sensitive data (secret)
as clear text.
return
with open(log_file_path, "r") as log_file:
for this_line in log_file.readlines()[lines * -1 :]:
Expand All @@ -291,7 +334,7 @@
executable = VFSProcessManager.find_vfs_launch_script()

command = (
f"sudo -E -u {self._os_user}"
f"{self._resolve_or_raise('sudo')} -E -u {self._os_user}"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Related but distinct from the shutdown-path comment: on the mount path, routing a missing findmnt into VFSExecutableMissingError converts what used to be a hard failure into a silent copy-fallback that races a live VFS process.

Sequence when sudo resolves but findmnt does not:

  1. asset_sync calls VFSProcessManager.find_vfs() (asset_sync.py:280 / :917), which succeeds — it uses shutil.which/$DEADLINE_VFS_INSTALL_PATH and knows nothing about findmnt.
  2. mount_vfs_from_manifests -> vfs_manager.start() (download.py:1258) reaches build_launch_command here, resolves sudo fine, and subprocess.Popen launches the VFS successfully (vfs.py:539).
  3. start then calls wait_for_mount -> is_mount -> _resolve_or_raise("findmnt"), which raises VFSExecutableMissingError. Nothing in start catches it — the try only wraps the Popen block — so it propagates out.
  4. asset_sync catches VFSExecutableMissingError and logs Virtual File System not found, falling back to COPIED, then runs download_files_from_manifests over the same merged_manifests_by_root.

The launched deadline_vfs process is still alive and mounting (or about to mount) those roots, and it was never recorded in the pid file — start raises before the pid-file write at vfs.py:568. So the copy-based download writes into roots that a live FUSE mount may claim, and cleanup_session -> kill_all_processes has no pid entry to clean up.

Before this change, subprocess.run(["findmnt", path]) raised FileNotFoundError, which is not VFSExecutableMissingError, so it propagated out of sync_inputs as a hard error rather than being absorbed by the fallback. Turning it into the fallback exception is only safe if the fallback is reached before anything is launched, which is not the case here.

Two things would close this independently of the findmnt question:

  • Resolve the commands start needs up front (before Popen), so a resolution failure cannot happen after a process exists.
  • Or have start kill/reap self._vfs_proc when anything after Popen fails, not just on VFSFailedToMountError.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is the sharpest of the three and I want to be clear I am not dismissing it. A silent copy-fallback that races a live VFS process is a worse outcome than a hard failure, and you are right that translating to VFSExecutableMissingError is what makes it silent.

Deferred on reachability: it needs sudo to resolve while findmnt does not, and both are in /usr/bin on every supported host. It is recorded as a real-but-unreachable finding rather than closed. If you would rather the mount path fail hard on a resolution failure, that is a one-line change to which exception _resolve_or_raise raises for that call site, and I will make it if you confirm that is the behaviour you want.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Scope question on the threat model, now that this line is the one being hardened: the string built here is handed to subprocess.Popen(args=start_command, shell=True, executable="/bin/bash") (vfs.py:558-565), and every field after sudo is interpolated unquoted:

f" {executable} {mount_point} -f --clienttype=deadline"
f" --bucket={self._asset_bucket}"
f" --manifest={self._manifest_path}"
...
command += f" --casprefix={self._cas_prefix}"
command += f" --cachedir={self._asset_cache_path}"

mount_point and _manifest_path derive from the session directory and the job's asset roots, and _cas_prefix/_asset_cache_path from queue/job configuration — i.e. values that come from the job being run rather than from this process. Any shell metacharacter in one of them (;, $(...), a space) is interpreted by /bin/bash, and the resulting command runs under the sudo this PR just pinned. That is a strictly wider vector than PATH resolution: CWE-426 requires influence over the search path, whereas this needs only one odd character in a path that is already carried through the job.

This is pre-existing, so it is fair to keep out of scope. But it does mean the hardening on this line closes the narrower of the two issues in the same expression, and the PR reads as though command construction here has been made safe. Either passing an argv list (which removes shell=True and the quoting question together — the resolved sudo path stays absolute, so the PATH property is preserved) or shlex.quote-ing the interpolated fields would close it; failing that, a note that shell quoting is knowingly separate would keep the boundary honest.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed on the substance, and I have taken the third option you offered: the note. Documented in 6f09326 on build_launch_command.

You are right that this is the wider of the two issues in that expression. Resolving sudo fixes how one binary is located; it does nothing about the fields interpolated unquoted after it, and those come from the job and its queue configuration rather than from this process. Needing one odd character in a path is a lower bar than influence over PATH.

Not fixing it here because both routes change launch mechanics rather than command resolution. The argv-list version is the one I would pick, since it removes shell=True and the quoting question together while keeping the resolved absolute sudo path, but it also changes how the environment and executable="/bin/bash" are handled, which is more than this PR should carry.

The part I did want to close is your last point, that the PR reads as though command construction here has been made safe. The note says plainly what is and is not addressed, so the boundary is on the record rather than implied.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

findmnt is pre-resolved in start() but sudo is not, and resolving it here changes an unresolvable-sudo host from a hard failure into a COPIED download into a 0o777 directory.

Order in start():

self._resolve_or_raise("findmnt")                        # vfs.py:591  (pre-resolved)
self.set_manifest_owner()
VFSProcessManager.create_mount_point(self._mount_point)  # vfs.py:593  -> os.chmod(mode=0o777)
start_command = self.build_launch_command(...)           # vfs.py:594  -> resolves sudo, may raise

create_mount_point (vfs.py:522-530) does os.makedirs then os.chmod(path=mount_point, mode=0o777). So when sudo is not in a trusted directory, _resolve_or_raise("sudo") on this line raises VFSExecutableMissingError after the asset root has been made world-writable — and asset_sync (asset_sync.py:292, :931) catches that exception and falls back to download_files_from_manifests over the same merged_manifests_by_root. The copy-based download then writes job input files into a 0o777 directory on a multi-tenant worker.

That is a behaviour change from the base commit, not just a relocation. With the bare "sudo" string, build_launch_command could not fail: Popen(shell=True) succeeded, bash reported sudo: command not found, wait_for_mount timed out, and start() raised VFSFailedToMountError — which no caller handles, so it surfaced as a hard error and never reached the COPIED path. Routing the missing-sudo case into VFSExecutableMissingError moves it into the fallback, which is the right type for the fallback but is now reached after mutating the filesystem.

The pre-resolve added at vfs.py:591 already establishes the pattern that fixes this — resolving sudo alongside findmnt, before create_mount_point, makes the whole resolve-then-mutate ordering hold, and the fallback would then run against an untouched root. test_start_resolves_findmnt_before_launching asserts mock_create_mount_point.assert_not_called() for findmnt; the same assertion for sudo currently would not pass.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Deferred. Correct that sudo resolves inside build_launch_command while findmnt is now pre-resolved, so the two missing-binary cases do not fail alike.

The asymmetry is smaller than the one that was fixed: both raise before Popen, so neither leaves a process running, which was the defect. Aligning them is a small change but it moves resolution out of the function that builds the command, and this PR has had enough churn in start(). Filed.

f" {executable} {mount_point} -f --clienttype=deadline"
f" --bucket={self._asset_bucket}"
f" --manifest={self._manifest_path}"
Expand Down Expand Up @@ -344,6 +387,20 @@
Determine where the VFS executable we'll be launching lives so we can
find the correct relative paths around it for LD_LIBRARY_PATH and config files
:return: Path to VFS executable

Note this deliberately does not use the trusted-directory resolver in
``_system_commands``, and the difference is worth understanding before
changing either one. That resolver exists for *system* commands, whose
locations are fixed and small in number. The VFS executable is a shipped
artifact whose location is a deployment choice: this function consults
``shutil.which``, then ``$DEADLINE_VFS_INSTALL_PATH``, then a path relative to
the working directory, because a deployment may legitimately put it in any of
them.

So the search path for this binary, and for the launch script found by
:meth:`find_vfs_launch_script`, remains wider than for ``sudo`` or
``findmnt``. Narrowing it is a separate question about how the VFS is
deployed, not about command resolution, and it is knowingly out of scope here.
"""
if VFSProcessManager.exe_path is not None:
log.info(f"Using saved path {VFSProcessManager.exe_path}")
Expand Down Expand Up @@ -417,7 +474,7 @@
"""
if VFSProcessManager.cwd_path is None:
exe_path = VFSProcessManager.find_vfs()
# Use cwd one folder up from bin

Check failure

Code scanning / CodeQL

Clear-text logging of sensitive information High

This expression logs
sensitive data (secret)
as clear text.
VFSProcessManager.cwd_path = os.path.normpath(
os.path.join(os.path.dirname(exe_path), "..")
)
Expand Down Expand Up @@ -447,7 +504,7 @@
log.error(f"Manifest not found at {self._manifest_path}")
return
if self._os_group is not None:
try:

Check failure

Code scanning / CodeQL

Clear-text logging of sensitive information High

This expression logs
sensitive data (secret)
as clear text.
shutil.chown(self._manifest_path, group=self._os_group)
os.chmod(self._manifest_path, DEADLINE_MANIFEST_GROUP_READ_PERMS)
except OSError as e:
Expand Down
Loading
Loading