Skip to content

kv-router: skip partial-prefix BlockStored events instead of exiting - #38

Draft
Thachnh wants to merge 1 commit into
prod-kv-indexersfrom
v41-partial-skip
Draft

kv-router: skip partial-prefix BlockStored events instead of exiting#38
Thachnh wants to merge 1 commit into
prod-kv-indexersfrom
v41-partial-skip

Conversation

@Thachnh

@Thachnh Thachnh commented Sep 10, 2026

Copy link
Copy Markdown

Problem

vLLM (from the DeepSeek-V4.1-Flash day-0 image, vllm/vllm-openai:deepseekv41-flash-0909) hashes prefixes every prefix_match_unit (hash_block_size, 32 for V4.1) tokens, which is finer than the cache block size (128). It publishes the prompt tail that ends inside a cache block as a BlockStored on the main MLA group whose block_size is the sub-block length (32/64/96).

The standalone indexer treats any main-attention block-size mismatch as a fatal --block-size misconfiguration and calls std::process::exit(1), so all three kv-indexer:* flavors for deepseek-ai/DeepSeek-V4.1-Flash crash-looped (engine block size 32 != configured --block-size 128). No indexer --block-size works: 128 dies on the partial events, 32 dies on the whole-block events.

Observed event mix on a V4.1 pod (150 s tap): groups 0-3 sliding_window_mla @32 (already filtered as non-main), group 4 mla_attention @128 (whole blocks) plus @32 partial entries.

Fix

convert_event: when block_size < kv_block_size treat the event as a partial-prefix entry and drop it (rate-limited warn) instead of erroring. A larger event block size stays fatal (that can only be a misconfig). Whole-block routing information is unaffected; only sub-block tail entries are ignored.

Unit test added; cargo test --lib zmq_wire::tests = 19 passed.

Same commit cherry-picked onto 24h-indexer as v41-partial-skip-h24 for the h24 image.

Refs DEE-640.

vLLM (>= the DeepSeek-V4.1 day-0 image) hashes prefixes every
prefix_match_unit (hash_block_size) tokens, which can be finer than the
cache block size, and publishes the prompt tail that ends inside a cache
block as a BlockStored whose block_size is the sub-block length (32/64/96
for a 128-token block). The standalone indexer treated any block_size
mismatch on a main-attention event as a fatal --block-size misconfig and
exited, so every V4.1 indexer crash-looped.

Treat a proper divisor of the configured block size as a partial-prefix
entry: drop it (rate-limited warn). Non-divisor mismatches stay fatal.
@Thachnh
Thachnh deployed to external_collaborator September 10, 2026 17:37 — with GitHub Actions Active
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant