Skip to content

fix(kv-router): skip offload-tier stores and map vLLM cpu mediums to host tier - #50

Merged
ovuruska merged 2 commits into
24h-indexerfrom
h24-skip-offload-block-size-mismatch
Oct 6, 2026
Merged

ovuruska merged 2 commits into
24h-indexerfrom
h24-skip-offload-block-size-mismatch

Conversation

@ovuruska

@ovuruska ovuruska commented Oct 6, 2026 •

Copy link
Copy Markdown

Summary

Two fixes for engines that run vLLM with a CPU offload connector, such as Kimi-K3 with Mooncake:

  1. Offload stores no longer crash the indexer. The standalone indexer exited on the first BlockStored whose block_size differed from --block-size ("Fatal KV event config error"). vLLM's offload connector emits lower-tier stores at its own block size, so such engines crashlooped the indexer whatever --block-size it was given. The normalizer now filters those events under a new reason, offload_block_size_mismatch, with its own rate-limited warning budget. A block-size mismatch on the device tier stays fatal.
  2. vLLM medium names map to the right tier. StorageTier::from_kv_medium was case-sensitive and did not know vLLM's own names, so "cpu" fell back to the device tier. Two consequences: a host-tier BlockRemoved evicted the same hash from the GPU tier, and same-size host stores landed in the GPU tree. It now matches case-insensitively and accepts CPU -> HostPinned and STORAGE -> Disk. This also changes lib/llm's KVBM tracker (from_vllm_medium), which previously returned None for "CPU".

Evidence (moonshotai/Kimi-K3, Dynamo agg, 3 workers)

30 s of the raw ZMQ stream on :5557, per worker:

group_idx kv_cache_spec_kind medium block_size
0, 1, 2 mamba GPU 1536 (already filtered as non-main)
3 mla_attention GPU 1536 (the main KV events)
3 (none) cpu 12288 (one 12288-token block = 8 x 1536)

With --block-size=1536 the indexer died on the cpu events; with 12288 it would die on the GPU ones. Nothing is lost by skipping them: those blocks were already stored, and indexed, when the GPU tier stored them.

The image from the first commit (kvtest-7470ac7ea3) has run Kimi-K3's h24 indexer in prod since 13:14 UTC 2026-10-06 without a restart. h24 ignores Removed events, so issue 2 does not affect it; it does affect reality and routing indexers on offloading engines.

Testing

  • cargo test -p dynamo-kv-router --lib: 553 passed.
  • New tests:
    • test_normalizer_skips_offload_tier_block_size_mismatch: cpu-tier stores at a different size are filtered.
    • test_normalizer_keeps_device_block_size_mismatch_fatal: a mismatch with no medium or with GPU still returns BlockSizeMismatch.
    • test_storage_tier_accepts_vllm_medium_names: cpu/CPU map to HostPinned, STORAGE to Disk, gpu to Device, unknown names to None.
    • test_normalizer_keeps_cpu_events_on_the_host_tier: cpu BlockStored and BlockRemoved land on HostPinned.
    • test_offload_skips_do_not_spend_the_shared_warning_budget: offload skips leave the shared warning counter at 0.
  • cargo fmt --check: clean for the touched code.
  • lib/llm was not rebuilt; from_kv_medium's signature is unchanged.

Rollout

Build with container/build-indexer-image.sh, push, then point only Kimi-K3's kv-indexer:h24 at the new tag. Other models' indexers stay on their current images.

🤖 Generated with Claude Code

vLLM's CPU offload connector emits BlockStored events for the lower tier at
its own block size (several GPU blocks per offloaded block; medium "cpu").
The block-size check treated them as a config error and the standalone
indexer exited, so any model with CPU offload (moonshotai/Kimi-K3: GPU 1536,
CPU 12288) crashlooped regardless of --block-size.

Filter such events in the normalizer as offload_block_size_mismatch instead.
The GPU tier already stored those blocks, so nothing is lost for the index;
a device-tier mismatch stays fatal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ovuruska
ovuruska deployed to external_collaborator October 6, 2026 13:09 — with GitHub Actions Active
StorageTier::from_kv_medium was case-sensitive and did not know vLLM's own
medium names, so "cpu" (sent by vLLM offload connectors) fell back to the
device tier: a host-tier BlockRemoved evicted the same hash from the GPU
tier, and same-size host stores landed in the GPU tree. Match
case-insensitively and accept CPU -> HostPinned, STORAGE -> Disk.

Also give the offload-skip warning its own budget, so CPU offload chatter
at startup no longer silences the shared "Block not published" warnings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ovuruska
ovuruska had a problem deploying to external_collaborator October 6, 2026 15:23 — with GitHub Actions Failure
@ovuruska ovuruska changed the title kv-router: skip offload-tier stores with a different block size fix(kv-router): skip offload-tier stores and map vLLM cpu mediums to host tier Oct 6, 2026
@github-actions github-actions Bot added the fix label Oct 6, 2026
@ovuruska
ovuruska merged commit 3b04e9e into 24h-indexer Oct 6, 2026
13 of 27 checks passed

This branch had an error being deployed

1 failed deployment
external_collaborator — 12e00b27 Deployed Oct 6, 2026 by ovuruska via ok-to-test #57
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant