Repository navigation
standalone-indexer: hash image blocks from tokens only - #43
Merged
Merged
Conversation
vLLM attaches each image's identifier to a stored block in extra_keys, and the shared ZMQ normalizer mixed it into the block's tokens hash. The standalone indexer's queriers hash plain token ids: deepapi's probe does, and so do the engine local-indexer TreeDumps served by /kv_recover. Every query therefore stopped matching at a conversation's first image block. Measured on frank/DeepSeek-V4.1-Flash with a live ZMQ capture: for chains whose first image fell in the window, a plain probe matched exactly up to the image block (217, 1098, 1569, 358 blocks) while an image-aware probe matched the whole chain (589, 2432, 3054, 6450). The indexer predicted 0.77 vs actual 0.96 on prompts over 500k tokens. Add ZmqEventNormalizer::with_plain_mm_hashing() and use it in the standalone listener. The library default is unchanged for Dynamo's own router, which queries with image info. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Author
|
h24-line companion: #44 (615 tests passed there). |
Shang-Pin
changed the base branch from
recover-hardening
to
v41-partial-skip
September 22, 2026 23:07
Shang-Pin
marked this pull request as ready for review
September 22, 2026 23:07
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
vLLM attaches each image's identifier to stored blocks in
extra_keys. The shared ZMQ normalizer mixes it into the block's tokens hash, but the standalone indexer's queriers hash plain token ids: deepapi's probe, and the engine local-indexer TreeDumps served by/kv_recover. Every query stops matching at a conversation's first image block, for all three flavors (reality, routing, h24).Measured on frank/DeepSeek-V4.1-Flash, 2026-09-22:
or_conv completelines: indexer 0.931 vs actual 0.923 under 100k tokens, but 0.769 vs 0.961 over 500k (long agent sessions with screenshots).Change
ZmqEventNormalizer::with_plain_mm_hashing()stripsblock_mm_infosbefore conversion.Trade-off: prompts sharing text but carrying different images at the same position index as the same prefix, so the indexer can over-report a hit.
Validation
cargo test -p dynamo-kv-router --features kube-discovery --lib: 608 passed, including the newtest_plain_mm_hashing_matches_token_only_probe.Stacked on #41.
🤖 Generated with Claude Code