Skip to content

kv-router: route vLLM "CPU" KV events to the HostPinned tier (#10368) - #47

Draft
Thachnh wants to merge 2 commits into
mm-aware-hashfrom
cpu-medium-host-tier
Draft

Thachnh wants to merge 2 commits into
mm-aware-hashfrom
cpu-medium-host-tier

Conversation

@Thachnh

@Thachnh Thachnh commented Sep 23, 2026

Copy link
Copy Markdown

Cherry-picks upstream ai-dynamo/dynamo ai-dynamo#10368 (d8a94ef28): map vLLM's medium="CPU" KV events to the HostPinned tier.

Why

vLLM's SimpleCPUOffloadConnector tags its KV events medium="CPU". Our fork branched from upstream on 2026-06-02, before ai-dynamo#10368, and only maps "CPU_PINNED"/"CPU_TIER1". So "CPU" fell through from_kv_medium_or_default to the Device tier:

  • a CPU-tier eviction deleted a block the GPU still held, so routing thought the pod had lost the prefix;
  • a block evicted from GPU but still in the CPU tier looked GPU-resident, so deepapi couldn't rank it below a real GPU hit.

Confirmed on a zai-org/GLM-5.3 prod pod (60 s of its ZMQ stream): GPU stored 7,884 / removed 7,871; CPU stored 8,594 / removed 8,594. The events are balanced and correctly tagged, so the engine is fine and the bug is in the indexer. Every deployed kvtest-* image has it. It affects every model that runs CPU offload with the indexer (GLM-5.3, DeepSeek-V4-Pro, Kimi). On GLM-5.3, DeepApi:kv_tokens reality_best (80.2%) sits below actual (82.4%) over 6 h.

deepapi already consumes tiers (kv_shards_from_indexer_response: gpu vs longest_matched → PRIMARY/SECONDARY), so no backend change is needed.

Changes

  • protocols.rs: "CPU" | "CPU_PINNED" | "CPU_TIER1" => HostPinned (upstream)
  • lower-tier indexers share the metrics handle, so HostPinned traffic is counted (upstream)
  • zmq_wire/tests.rs: conflict resolved by keeping all our tests and adding test_convert_event_cpu_medium_lands_on_host_tier. Upstream's zero-block-size CPU test is dropped because our branch already skips those as partial-prefix entries (90409c7f8).

Validation

  • cargo test -p dynamo-kv-router: 550 passed, 0 failed
  • image localhost:30500/dynamo-indexer:kvtest-7d461346ca built and pushed
  • the same commit on the h24 line applied cleanly (companion PR)

Not done

Refs DEE-720, DEE-722

https://claude.ai/code/session_01575gLgRp2Xk4WLR9LsiMu6

…mo#10368)

Cherry-pick of upstream d8a94ef. vLLM's SimpleCPUOffloadConnector tags
its events medium="CPU"; our fork only mapped "CPU_PINNED"/"CPU_TIER1",
so "CPU" fell back to the Device tier. CPU-tier evictions then removed
blocks the GPU still held, and CPU-only blocks looked GPU-resident.
Verified on a zai-org/GLM-5.3 pod: CPU stores/removes are balanced and
tagged "CPU". Conflict in zmq_wire/tests.rs resolved by keeping our tests
and adding one for the medium -> tier mapping (upstream's zero-block-size
CPU test dropped: our branch skips those as partial-prefix entries).

(cherry picked from commit d8a94ef)

Claude-Session: https://claude.ai/code/session_01575gLgRp2Xk4WLR9LsiMu6
@Thachnh
Thachnh deployed to external_collaborator September 23, 2026 22:16 — with GitHub Actions Active
…eeds them

With the "CPU" medium routed to the HostPinned tier (ai-dynamo#10368), a device
store whose parent the device tree only knows from the host tier is
rejected with ParentBlockNotFound, and every later block of that
sequence is rejected too. On zai-org/GLM-5.3 (SimpleCPUOffloadConnector)
388/388 rejected stores in a sample had a host-only parent, ~100-150/min
per offload pod after the image swap.

The engine can only cache a device block whose parent is resident on the
device, so a device store proves its ancestor chain is there. Each
listener now keeps a TierBridge: per (worker, dp_rank) sets of device and
host block hashes (host with parent links), updated in event order from
live and recovered events and reset on TreeDump. When a device store's
parent is host-only, the bridge walks the host chain up to the nearest
device block (or the root) and the listener applies that chain as a
device store before the original event. A broken chain is left alone.

Claude-Session: https://claude.ai/code/session_01575gLgRp2Xk4WLR9LsiMu6
@Thachnh
Thachnh had a problem deploying to external_collaborator September 24, 2026 00:55 — with GitHub Actions Failure
@Thachnh

Thachnh commented Sep 24, 2026

Copy link
Copy Markdown
Author

Update: added b3a2feceac (standalone-indexer TierBridge: promote host-only ancestors when a device store needs them) → image kvtest-11cd973221, live on zai-org/GLM-5.3 reality + routing since 2026-09-24 00:56 UTC (routing on prompt-cache-key throughout).

The bridge does not fix the problem. ParentBlockNotFound is still 150–300/min after warm-up. In the /dump, 458 of 568 rejected stores (81%) have an ancestor in neither tier, so no indexer-side repair can rebuild them. The likely cause is engine-side (the TreeDump / local indexer in our vLLM kv-events patch). That's tracked in DEE-725.

Recommendation for this PR: keep 7d461346ca (ai-dynamo#10368 CPU→HostPinned), since it's the correct tier mapping. Consider reverting b3a2feceac unless DEE-725 shows it helps for the remaining ~19%.

metrics::tests::register_and_encode fails with --features indexer-runtime on the base branch too (pre-existing).

This branch had an error being deployed

1 failed deployment
external_collaborator — 11cd9732 Deployed Sep 24, 2026 by Thachnh via ok-to-test #53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants