Skip to content

standalone-indexer: subscribe to every data-parallel rank of a discovered pod - #39

Merged
Shang-Pin merged 1 commit into
v41-partial-skipfrom
dp-rank-watch
Sep 17, 2026
Merged

Shang-Pin merged 1 commit into
v41-partial-skipfrom
dp-rank-watch

Conversation

@Shang-Pin

Copy link
Copy Markdown

Why

Pod discovery registers each engine pod once, as dp_rank 0 on tcp://<ip>:<zmq_port>. vLLM offsets both the KV-event ZMQ port and the /kv_recover port by data_parallel_rank, so on a --data-parallel-size N engine, ranks 1..N-1 publish to ports nobody subscribes to. Everything they cache is invisible to /query while the engine still hits it.

Measured on deepseek-ai/DeepSeek-V4.1-Flash (DP=2 in the main-0915-ds4-patches image): every prefill the scheduler placed on rank 1, about half of them, showed 0 in the standalone indexer and in rank 0's /kv_recover dump, and all of its blocks in rank 1's dump. deepapi's KV ladder read actual 0.78 > h24 0.65 > perfect 0.63 > reality 0.59 (engine's own hit ratio 0.782), i.e. the indexer under-reports the cache by about a third and KV-aware routing is blind to half of every pod.

What

  • New --watch-dp-size N (default 1, so single-rank engines are unchanged) carried as KubeDiscoveryConfig::dp_size.
  • The watcher registers one listener per rank under the pod's instance: dp_rank r, tcp://<ip>:<zmq_port + r>, http://<ip>:<recover_port + r> (rank_endpoints).
  • register() rejects a rank that is already present, so a pod whose registration fails part-way is deregistered before the retry instead of being left half-subscribed forever.
  • Unit tests for the per-rank endpoints; docs.md section.

Rollout

i models -n deepseek-ai/DeepSeek-V4.1-Flash services update kv-indexer:{routing,reality,h24} \
    --container-image localhost:30500/dynamo-indexer:kvtest-<sha> --extra-arg=--watch-dp-size=2
i models -n deepseek-ai/DeepSeek-V4.1-Flash services deploy

Verify: /workers lists two listeners per pod (ports 5557/5558); a fresh prefill matches in the indexer regardless of the rank that ran it; the ladder returns to h24 >= perfect >= actual.

Follow-up (backend): derive --watch-dp-size from the model config's --data-parallel-size at services deploy so the next DP model does not repeat this.

🤖 Generated with Claude Code

…ered pod

Pod discovery registered each engine pod once, as dp_rank 0 on
tcp://<ip>:<zmq_port>. vLLM offsets the KV-event ZMQ port and the
/kv_recover port by data_parallel_rank, so on a --data-parallel-size N
engine ranks 1..N-1 publish to ports nobody subscribes to and everything
they cache is invisible to /query while the engine still hits it.

Measured on deepseek-ai/DeepSeek-V4.1-Flash (DP=2): every prefill the
scheduler placed on rank 1 (about half) showed 0 in the standalone
indexer and in rank 0's /kv_recover dump, and all its blocks in rank 1's
dump; the KV ladder read actual 0.78 > h24 0.65 > perfect 0.63 >
reality 0.59.

Add --watch-dp-size (default 1, so single-rank engines are unchanged).
The watcher registers one listener per rank under the pod's instance:
dp_rank r, tcp://<ip>:<zmq_port + r>, http://<ip>:<recover_port + r>.
register() rejects a rank that is already present, so a pod whose
registration fails part-way is deregistered before the retry instead of
being left half-subscribed forever.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Shang-Pin
Shang-Pin deployed to external_collaborator September 16, 2026 23:25 — with GitHub Actions Active
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 16, 2026
@Shang-Pin

Copy link
Copy Markdown
Author

Built and pushed localhost:30500/dynamo-indexer:kvtest-bb6dcfd5c5 from this branch (container/build-indexer-image.sh). cargo test -p dynamo-kv-router --features kube-discovery pod_watcher: 11 passed (4 new), no warnings in the touched files. Rollout to deepseek-ai/DeepSeek-V4.1-Flash's three indexers pending.

@Shang-Pin

Copy link
Copy Markdown
Author

Rolled out to deepseek-ai/DeepSeek-V4.1-Flash kv-indexer:routing and kv-indexer:reality (kvtest-bb6dcfd5c5, --watch-dp-size=2). /workers now shows listeners on 5557 and 5558 for all six pods, all active. Fresh 50k-token prefills on one pod: 3/3 matched in the reality indexer (1/3 before). KV ladder over the first 10 min: reality_best 0.80, perfect 0.79, actual 0.75 (was reality 0.59 / perfect 0.63 / actual 0.78); h24 stays 0.59 until #40 lands, since its image comes from the h24 branch.

@Shang-Pin
Shang-Pin marked this pull request as ready for review September 17, 2026 21:14
@Shang-Pin
Shang-Pin merged commit 87dd5a2 into v41-partial-skip Sep 17, 2026
14 of 26 checks passed

This branch was successfully deployed

1 active deployment
external_collaborator — bb6dcfd5 Deployed Sep 16, 2026 by Shang-Pin via ok-to-test #37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant