Repository navigation
standalone-indexer: subscribe to every data-parallel rank of a discovered pod - #39
Conversation
…ered pod Pod discovery registered each engine pod once, as dp_rank 0 on tcp://<ip>:<zmq_port>. vLLM offsets the KV-event ZMQ port and the /kv_recover port by data_parallel_rank, so on a --data-parallel-size N engine ranks 1..N-1 publish to ports nobody subscribes to and everything they cache is invisible to /query while the engine still hits it. Measured on deepseek-ai/DeepSeek-V4.1-Flash (DP=2): every prefill the scheduler placed on rank 1 (about half) showed 0 in the standalone indexer and in rank 0's /kv_recover dump, and all its blocks in rank 1's dump; the KV ladder read actual 0.78 > h24 0.65 > perfect 0.63 > reality 0.59. Add --watch-dp-size (default 1, so single-rank engines are unchanged). The watcher registers one listener per rank under the pod's instance: dp_rank r, tcp://<ip>:<zmq_port + r>, http://<ip>:<recover_port + r>. register() rejects a rank that is already present, so a pod whose registration fails part-way is deregistered before the retry instead of being left half-subscribed forever. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
|
Built and pushed |
|
Rolled out to deepseek-ai/DeepSeek-V4.1-Flash kv-indexer:routing and kv-indexer:reality ( |
Why
Pod discovery registers each engine pod once, as
dp_rank 0ontcp://<ip>:<zmq_port>. vLLM offsets both the KV-event ZMQ port and the/kv_recoverport bydata_parallel_rank, so on a--data-parallel-size Nengine, ranks 1..N-1 publish to ports nobody subscribes to. Everything they cache is invisible to/querywhile the engine still hits it.Measured on
deepseek-ai/DeepSeek-V4.1-Flash(DP=2 in themain-0915-ds4-patchesimage): every prefill the scheduler placed on rank 1, about half of them, showed 0 in the standalone indexer and in rank 0's/kv_recoverdump, and all of its blocks in rank 1's dump. deepapi's KV ladder readactual 0.78 > h24 0.65 > perfect 0.63 > reality 0.59(engine's own hit ratio 0.782), i.e. the indexer under-reports the cache by about a third and KV-aware routing is blind to half of every pod.What
--watch-dp-size N(default 1, so single-rank engines are unchanged) carried asKubeDiscoveryConfig::dp_size.dp_rank r,tcp://<ip>:<zmq_port + r>,http://<ip>:<recover_port + r>(rank_endpoints).register()rejects a rank that is already present, so a pod whose registration fails part-way is deregistered before the retry instead of being left half-subscribed forever.docs.mdsection.Rollout
Verify:
/workerslists twolistenersper pod (ports 5557/5558); a fresh prefill matches in the indexer regardless of the rank that ran it; the ladder returns toh24 >= perfect >= actual.Follow-up (backend): derive
--watch-dp-sizefrom the model config's--data-parallel-sizeatservices deployso the next DP model does not repeat this.🤖 Generated with Claude Code