Repository navigation
Conversation
…mo#10368) Cherry-pick of upstream d8a94ef. vLLM's SimpleCPUOffloadConnector tags its events medium="CPU"; our fork only mapped "CPU_PINNED"/"CPU_TIER1", so "CPU" fell back to the Device tier. CPU-tier evictions then removed blocks the GPU still held, and CPU-only blocks looked GPU-resident. Verified on a zai-org/GLM-5.3 pod: CPU stores/removes are balanced and tagged "CPU". Conflict in zmq_wire/tests.rs resolved by keeping our tests and adding one for the medium -> tier mapping (upstream's zero-block-size CPU test dropped: our branch skips those as partial-prefix entries). (cherry picked from commit d8a94ef) Claude-Session: https://claude.ai/code/session_01575gLgRp2Xk4WLR9LsiMu6
…eeds them With the "CPU" medium routed to the HostPinned tier (ai-dynamo#10368), a device store whose parent the device tree only knows from the host tier is rejected with ParentBlockNotFound, and every later block of that sequence is rejected too. On zai-org/GLM-5.3 (SimpleCPUOffloadConnector) 388/388 rejected stores in a sample had a host-only parent, ~100-150/min per offload pod after the image swap. The engine can only cache a device block whose parent is resident on the device, so a device store proves its ancestor chain is there. Each listener now keeps a TierBridge: per (worker, dp_rank) sets of device and host block hashes (host with parent links), updated in event order from live and recovered events and reset on TreeDump. When a device store's parent is host-only, the bridge walks the host chain up to the nearest device block (or the root) and the listener applies that chain as a device store before the original event. A broken chain is left alone. Claude-Session: https://claude.ai/code/session_01575gLgRp2Xk4WLR9LsiMu6
|
Update: added The bridge does not fix the problem. Recommendation for this PR: keep
|
Cherry-picks upstream ai-dynamo/dynamo ai-dynamo#10368 (
d8a94ef28): map vLLM'smedium="CPU"KV events to theHostPinnedtier.Why
vLLM's
SimpleCPUOffloadConnectortags its KV eventsmedium="CPU". Our fork branched from upstream on 2026-06-02, before ai-dynamo#10368, and only maps"CPU_PINNED"/"CPU_TIER1". So"CPU"fell throughfrom_kv_medium_or_defaultto the Device tier:Confirmed on a
zai-org/GLM-5.3prod pod (60 s of its ZMQ stream): GPU stored 7,884 / removed 7,871; CPU stored 8,594 / removed 8,594. The events are balanced and correctly tagged, so the engine is fine and the bug is in the indexer. Every deployedkvtest-*image has it. It affects every model that runs CPU offload with the indexer (GLM-5.3, DeepSeek-V4-Pro, Kimi). On GLM-5.3,DeepApi:kv_tokensreality_best (80.2%) sits below actual (82.4%) over 6 h.deepapi already consumes tiers (
kv_shards_from_indexer_response:gpuvslongest_matched→ PRIMARY/SECONDARY), so no backend change is needed.Changes
protocols.rs:"CPU" | "CPU_PINNED" | "CPU_TIER1" => HostPinned(upstream)zmq_wire/tests.rs: conflict resolved by keeping all our tests and addingtest_convert_event_cpu_medium_lands_on_host_tier. Upstream's zero-block-size CPU test is dropped because our branch already skips those as partial-prefix entries (90409c7f8).Validation
cargo test -p dynamo-kv-router: 550 passed, 0 failedlocalhost:30500/dynamo-indexer:kvtest-7d461346cabuilt and pushedNot done
kv-indexer:realitywipes a model's routing tree for ~15 min, so the swap is scheduled separately.Refs DEE-720, DEE-722
https://claude.ai/code/session_01575gLgRp2Xk4WLR9LsiMu6