v0.7.1 still deadlocks in fuse_reverse_inval_inode on remote-deletion of files with open handles
Summary
hf-mount-fuse:v0.7.1 (which includes #192) still wedges with the same inval_inode
zombie+D-state signature #192 was meant to fix. On 2026-06-19, 3 of 4 stuck-Terminating
Spaces pods across sp-wp-aws-us-east-1-prod-111/112 were on v0.7.1 (the 4th on v0.6.2),
each with a wedged FUSE connection (waiting>0) and an unreapable zombie sidecar.
The key difference from the #192 scenario: this deadlock happens at runtime, not on shutdown.
#192 disarms the invalidator closures on SIGTERM; here the deadlock is already in place
minutes before the pod is ever asked to terminate, so the shutdown-path disarm never applies.
Environment
- hf-mount-fuse v0.7.1, hf-csi-driver v0.10.2
- Kernel FUSE (in-tree
fuse.ko), AL2023 nodes
- Reproduced on
r-radienlin-vid-…qcvc2 (prod-112, node ip-10-112-69-68), connection 650
(waiting=6, max_background=64). Same shape on prod-111 connections 311 (waiting=13) and
2896 (waiting=12).
Trigger (from sidecar logs)
At 09:49:41, ~4 min before the pod's deletionTimestamp (09:54:19), the poll loop detected a
batch of remote deletions of files the local app still had open handles on:
poll: Remote deletion detected: downloads/Asian/<file>.mp4
poll: Remote deletion of ino=32: unlinked path, kept orphan (open handles)
poll: Remote deletion of ino=27: unlinked path, kept orphan (open handles)
poll: Remote deletion of ino=28: unlinked path, kept orphan (open handles)
The daemon then issued notify_inval_inode for those inodes, which deadlocked against the
page/inode locks held by the app's in-flight read()/open() on the same inodes.
Kernel evidence (live, hung_task)
Daemon thread wedged in the kernel processing its own reverse invalidation (>860s):
INFO: task tokio-rt-worker:NNNN blocked for more than 860 seconds.
fuse_reverse_inval_inode+0x93/0xc0 [fuse]
fuse_notify+0x2db/0x520 [fuse]
fuse_dev_do_write+0x2b7/0x4e0 [fuse]
fuse_dev_write+0x53/0x80 [fuse]
App threads then block forever waiting on the daemon:
uvicorn (D) wchan=request_wait_answer
request_wait_answer → __fuse_simple_request → fuse_lookup_name → fuse_lookup
→ fuse_atomic_open → lookup_open → path_openat → __x64_sys_openat # open() on the mount
ffmpeg/ffprobe ×3 (D) wchan=folio_wait_bit_common
folio_wait_bit_common → filemap_update_page → filemap_get_pages → filemap_read
→ vfs_read → ksys_read # read() on the mount
Sidecar process is an unreapable zombie:
PID STAT CMD
NNN Zsl hf-mount-fuse-s <defunct>
Sidecar's own shutdown log corroborates (graceful flush ok, but unmount blocked):
Received shutdown signal, flushing 2 mount(s)
Shutting down VFS, flushing pending writes...
Flush loop finished, VFS shut down.
ERROR hf_mount::fuse: Failed to unmount "/tmp"
WARN hf_mount::fuse: Unmount failed, deferring to destroy() shutdown path
INFO hf_mount_fuse_sidecar: Flush complete, exiting # logs "exiting" but never reaps
Why #192 / hf-csi-driver#47 don't cover it
Proposed direction
- Make
notify_inval_inode non-blocking / abortable so a reverse invalidation can't deadlock
against an in-flight request holding the same page/inode lock (the runtime path, not just
shutdown).
- Operationally, a node-side sweep that aborts FUSE connections whose daemon is a zombie is the
version-independent escape hatch (see remediation below).
Operational remediation (works today)
echo 1 > /sys/fs/fuse/connections/<minor>/abort for the uid-matched minor → app unblocks,
kubelet reaps the pod in ~25s. Must target only the exact pod's minor, never a live neighbor's.
v0.7.1 still deadlocks in
fuse_reverse_inval_inodeon remote-deletion of files with open handlesSummary
hf-mount-fuse:v0.7.1(which includes #192) still wedges with the sameinval_inodezombie+D-state signature #192 was meant to fix. On 2026-06-19, 3 of 4 stuck-
TerminatingSpaces pods across
sp-wp-aws-us-east-1-prod-111/112were on v0.7.1 (the 4th on v0.6.2),each with a wedged FUSE connection (
waiting>0) and an unreapable zombie sidecar.The key difference from the #192 scenario: this deadlock happens at runtime, not on shutdown.
#192 disarms the invalidator closures on SIGTERM; here the deadlock is already in place
minutes before the pod is ever asked to terminate, so the shutdown-path disarm never applies.
Environment
fuse.ko), AL2023 nodesr-radienlin-vid-…qcvc2(prod-112, nodeip-10-112-69-68), connection 650(
waiting=6,max_background=64). Same shape on prod-111 connections 311 (waiting=13) and2896 (
waiting=12).Trigger (from sidecar logs)
At
09:49:41, ~4 min before the pod'sdeletionTimestamp(09:54:19), the poll loop detected abatch of remote deletions of files the local app still had open handles on:
The daemon then issued
notify_inval_inodefor those inodes, which deadlocked against thepage/inode locks held by the app's in-flight
read()/open()on the same inodes.Kernel evidence (live, hung_task)
Daemon thread wedged in the kernel processing its own reverse invalidation (>860s):
App threads then block forever waiting on the daemon:
Sidecar process is an unreapable zombie:
Sidecar's own shutdown log corroborates (graceful flush ok, but unmount blocked):
Why #192 / hf-csi-driver#47 don't cover it
operation; by the time SIGTERM arrives the daemon thread is already stuck in-kernel and the
disarm is moot.
NodeUnpublishVolumeas a backstop — but thepod never reaped, so either NodeUnpublish wasn't reached, or it doesn't cover the second
(
/tmp) mount. Worth confirming which mount(s) refactor: daemon spawns backends as subprocess #47's abort targets.Proposed direction
notify_inval_inodenon-blocking / abortable so a reverse invalidation can't deadlockagainst an in-flight request holding the same page/inode lock (the runtime path, not just
shutdown).
version-independent escape hatch (see remediation below).
Operational remediation (works today)
echo 1 > /sys/fs/fuse/connections/<minor>/abortfor the uid-matched minor → app unblocks,kubelet reaps the pod in ~25s. Must target only the exact pod's minor, never a live neighbor's.