Repository navigation
fix(kv-router): skip offload-tier stores and map vLLM cpu mediums to host tier - #50
Merged
Merged
Conversation
vLLM's CPU offload connector emits BlockStored events for the lower tier at its own block size (several GPU blocks per offloaded block; medium "cpu"). The block-size check treated them as a config error and the standalone indexer exited, so any model with CPU offload (moonshotai/Kimi-K3: GPU 1536, CPU 12288) crashlooped regardless of --block-size. Filter such events in the normalizer as offload_block_size_mismatch instead. The GPU tier already stored those blocks, so nothing is lost for the index; a device-tier mismatch stays fatal. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
StorageTier::from_kv_medium was case-sensitive and did not know vLLM's own medium names, so "cpu" (sent by vLLM offload connectors) fell back to the device tier: a host-tier BlockRemoved evicted the same hash from the GPU tier, and same-size host stores landed in the GPU tree. Match case-insensitively and accept CPU -> HostPinned, STORAGE -> Disk. Also give the offload-skip warning its own budget, so CPU offload chatter at startup no longer silences the shared "Block not published" warnings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ovuruska
had a problem deploying
to
external_collaborator
October 6, 2026 15:23 — with
GitHub Actions
Failure
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two fixes for engines that run vLLM with a CPU offload connector, such as Kimi-K3 with Mooncake:
BlockStoredwhoseblock_sizediffered from--block-size("Fatal KV event config error"). vLLM's offload connector emits lower-tier stores at its own block size, so such engines crashlooped the indexer whatever--block-sizeit was given. The normalizer now filters those events under a new reason,offload_block_size_mismatch, with its own rate-limited warning budget. A block-size mismatch on the device tier stays fatal.StorageTier::from_kv_mediumwas case-sensitive and did not know vLLM's own names, so"cpu"fell back to the device tier. Two consequences: a host-tierBlockRemovedevicted the same hash from the GPU tier, and same-size host stores landed in the GPU tree. It now matches case-insensitively and acceptsCPU->HostPinnedandSTORAGE->Disk. This also changeslib/llm's KVBM tracker (from_vllm_medium), which previously returnedNonefor"CPU".Evidence (moonshotai/Kimi-K3, Dynamo agg, 3 workers)
30 s of the raw ZMQ stream on
:5557, per worker:mambamla_attentioncpuWith
--block-size=1536the indexer died on the cpu events; with 12288 it would die on the GPU ones. Nothing is lost by skipping them: those blocks were already stored, and indexed, when the GPU tier stored them.The image from the first commit (
kvtest-7470ac7ea3) has run Kimi-K3's h24 indexer in prod since 13:14 UTC 2026-10-06 without a restart. h24 ignoresRemovedevents, so issue 2 does not affect it; it does affect reality and routing indexers on offloading engines.Testing
cargo test -p dynamo-kv-router --lib: 553 passed.test_normalizer_skips_offload_tier_block_size_mismatch: cpu-tier stores at a different size are filtered.test_normalizer_keeps_device_block_size_mismatch_fatal: a mismatch with no medium or withGPUstill returnsBlockSizeMismatch.test_storage_tier_accepts_vllm_medium_names:cpu/CPUmap toHostPinned,STORAGEtoDisk,gputoDevice, unknown names toNone.test_normalizer_keeps_cpu_events_on_the_host_tier: cpuBlockStoredandBlockRemovedland onHostPinned.test_offload_skips_do_not_spend_the_shared_warning_budget: offload skips leave the shared warning counter at 0.cargo fmt --check: clean for the touched code.lib/llmwas not rebuilt;from_kv_medium's signature is unchanged.Rollout
Build with
container/build-indexer-image.sh, push, then point only Kimi-K3'skv-indexer:h24at the new tag. Other models' indexers stay on their current images.🤖 Generated with Claude Code