Repository navigation
Conversation
…i-dynamo#11891) Signed-off-by: jthomson04 <64760228+jthomson04@users.noreply.github.com> (cherry picked from commit 6b78d18)
…amo#14878) Signed-off-by: jthomson04 <jwillthomson19@gmail.com> (cherry picked from commit 8f95101)
Signed-off-by: PeaBrane <yanrpei@gmail.com> (cherry picked from commit 6c66d4b)
Carries forward DeepInfra's indexer image recipe: a two-stage Ubuntu 24.04 build that compiles the dynamo._core wheel with kv-indexer-metrics and installs only that wheel into a slim runtime. Ports cd18634 (image + build script), ef5b033 (release build) and f4b4f1d (touch archived sources so a sibling branch's newer artifact in the shared cargo cache is never reused). Header comment now lists the three service flavors the one image serves. Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
A listener blocked in gap recovery can let its SUB pipe reach the default RCVHWM of 1000; with heartbeats on, libzmq 4.3.4 then aborts on `Assertion failed: _input_stopped` (zeromq/libzmq#3596, ai-dynamo#3937). Seen in prod as an h24 indexer crash loop while 66 startup TreeDumps applied. Only the indexer's connect_sub_socket changes; the shared receive config used by replica sync, selection and the slot tracker keeps its default. Ports ae7403c. Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
…x stores The shared normalizer drops every block whose size differs from the configured one, so an indexer started with the wrong --block-size stays empty while its stream looks alive. The standalone listener now checks each preprocessed BlockStored first: - a store of blocks larger than configured is a misconfiguration: log and exit so k8s surfaces a CrashLoopBackOff; - a store smaller than configured is a vLLM partial-prefix entry (prefix_match_unit < cache block, e.g. 32/64/96 of 128 on V4.1): drop it, logging the first three. Ports 959e440 and 90409c7. Unlike the originals this lives in the standalone listener (new block_size.rs) and leaves convert_event's signature and the lib/llm publisher untouched; the publisher half of 959e440 is Dynamo's own worker-side path, which we don't run. Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
…gate and retry
Our engines serve their local indexer at GET /kv_recover (HTTP) rather
than on a ZMQ ROUTER replay socket. A listener registered with a
recover_endpoint now recovers a sequence gap by fetching a
WorkerKvQueryResponse for [start, end) and applying it:
- Events: apply all under the consumer-assigned (worker, dp_rank),
resume at last_event_id;
- TreeDump (whole-rank scope): remove_worker_dp_rank, apply all, resume at
last_event_id; a domain-scoped dump applies without the reset;
- TooNew / InvalidRange / Error / TreeDumpFailed: no-op.
Without a recover_endpoint the upstream ZMQ replay path is unchanged.
Downloads share one reqwest client and a process-wide gate
(--recover-concurrency, default 8), with a total timeout of
--recover-timeout-secs (default 120) and 3 attempts with 2/4 s backoff
before the gap is given up ("kv_recover request failed; giving up").
v1.5.0 has no HTTP recovery at all (its only HTTP client is peer /dump
recovery, hard-coded to 10 s).
The settings are injected through WorkerRegistry::with_kv_recover rather
than process globals, and /register takes recover_endpoint as a new field
next to replay_endpoint instead of aliasing it. register() keeps
upstream's signature; register_with_extras() carries our fields.
Tests: a /kv_recover response captured from a prod GLM-5.3 engine
deserializes and plans under the consumer identity; retry until success;
give up after the last attempt; per-attempt timeout.
Ports ab35101 and 37e1b1a.
Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
With --watch-namespace plus --watch-model-name (or a raw --watch-label), the indexer watches the model's engine pods and keeps the registry in sync: a Ready pod is registered, a deleted / unready pod deregistered. kube::runtime's reflector re-lists on 410 Gone, and every change reconciles the desired Ready set against what is registered, so a missed delete is corrected on the next re-list. - The selector is di/model_name=<sanitized model name>, the label the backend stamps on every engine pod across engine_hash variants. - Each pod registers under routing_group = its engine_hash label (the configured group if unlabeled), so hash-incompatible engine generations never share a tree during a migration; a changed label re-registers the pod under its new group. - --watch-dp-size N registers one listener per data-parallel rank: dp_rank r on tcp://<ip>:<zmq_port + r>, recovering from http://<ip>:<recover_port + r>. A pod whose registration fails part-way is deregistered before the retry. - Configuring discovery in a build without the kube-discovery feature is now an error instead of a warning. Discovery config and the label helpers (discovery.rs) are always compiled; the kube client lives only in the feature-gated pod_watcher.rs. The python bindings' kv-indexer feature enables kube-discovery. Ports cd18634 (watcher), the model-name watch of d14e05d, the engine_hash grouping of b662110, and bb6dcfd (DP ranks). Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
Workers are keyed by instance_id (xxh3 of the pod name), which is one way. deepapi maps /query_by_hash matches back to shards by pod name (kv_shards_from_indexer_response), so keep the raw pod name on the worker record when the pod watcher or a /register caller provides it and surface it in /workers entries and in each /query `instances` entry. v1.5.0's shared MooncakeOverlapSummary has no pod_name; the indexer server wraps it in a flattened InstanceMatch rather than changing the shared type. Purely additive to the response. Adds deepapi_contract_tests.rs: the /query_by_hash request/response shape deepapi depends on. Ports b9a2249. Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
The pod watcher gives each engine_hash regime its own tree (routing group). A /query or /query_by_hash that names no routing_group now fans out over every tree of model_name and merges the per-tree responses: worker sets are disjoint, so scores and instances union; frequencies sum element-wise. A probe hashed in one regime scores 0 against the others, so routing converges on pods that can reuse the cache during image migrations. A request that names a routing_group still reads only that tree, as upstream does; the legacy tenant_id stays ignored. deepapi sends only tenant_id "default", so it gets the fan-out. A failing tree is skipped while any tree answers (500 only if all fail); 404 only when the model has no tree. The fan-out lives in model_query.rs and replaces upstream's single-tree run_tiered_query. Upstream's query_by_hash_isolates_routing_groups_and_ignores_legacy_tenant_id is renamed and its no-routing_group case now expects the fan-out. New tests: /query_by_hash across two groups excluding another model, /query by tokens, and deepapi's golden block hashes (plain and LoRA-seeded) from deepinfra/utils.kv_block_hashes. Ports the query half of b662110. Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
…mory pressure The kv-indexer:routing flavor measures "infinite memory, per-shard" hit rates: engine Removed events are parked in a per-listener buffer instead of applied (Cleared is dropped), and a store of a parked block cancels its eviction. Every 15 s, once memory use crosses --evict-memory-threshold (default 0.75) of the cgroup limit, evictions older than --evict-retention-secs (default 1800) are replayed into the tree as ordinary Removed events, so the pod degrades toward reality instead of OOMing. A per-minute status line logs memory fraction, pending and queued entries, and listener count. A replayed eviction is applied only if it is still the block's latest; a /kv_recover TreeDump that resets the rank clears the rank's buffer. The flag is a RegistryOptions field (no process globals): only listeners of a keep-evictions registry carry a buffer. Ports the keep-evictions half of d14e05d and c9eabbb; supersedes 6d16bb8 (--ignore-evictions / --merge-shards). Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
…s and events Off by default. When on, the indexer logs one kv_audit line per /query and /query_by_hash (model, routing groups queried, status, probe block hashes, full JSON response) and one per engine event it ingests (STORE / EVICT / CLEAR with worker, dp_rank, tier, event id, seq, and the token and sequence block hashes), tagged source=live|replay|recover and written before the --keep-evictions filter so it shows what the engine published. Isolate with RUST_LOG=kv_audit=info. Upstream's --access-log (ai-dynamo#10700) records requests but not hashes, responses or events. Ports 5e5f3b8. Not ported: the non-blocking stdout writer from d14e05d's CLI init; no prod deployment passes --enable-logging. Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
A new Indexer::H24 backend answers "how many prefix blocks of this query were stored by ANY worker within the retention horizon", ignoring evictions, clears and worker identity. Two flat maps replace the radix tree: engine sequence hash -> chained prefix hash (so stores naming their parent by engine hash keep chaining) and prefix hash -> last_stored / last_touched. Matches are reported under a synthetic worker 0, so the /query_by_hash response shape is unchanged. - Matched-block ages feed dynamo_kvindexer_h24_age_seconds, whose bucket CDF over queried blocks is the hit-rate-vs-retention curve; queried blocks are counted once per request, not per tree. - An expiry sweep every 10 min drops entries older than --h24-horizon-secs (default 172800), and tightens to half the horizon above 92% of the cgroup memory limit. - Every routing group of a model collapses into one h24 tree, so a block matched in several engine_hash regimes is counted once (the 120% hit-rate incident). - /dump returns nothing in h24 mode; --h24 and --keep-evictions are mutually exclusive. A runtime flag of the one image rather than a separate -h24 branch: every live h24 deployment already passes --h24. Ports 73474e5 and 247ac26. Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
Covers, through ListenerLoop with wire-format msgpack batches: - "CPU" medium stores land on the HostPinned tier, not the device tree (upstream ai-dynamo#10368; the fork this replaces filed them as device); - sub-block partial-prefix stores are skipped without exiting; - --keep-evictions parks an eviction and a re-store cancels it; - a sequence gap is filled from /kv_recover: a TreeDump replaces the rank's state under the consumer identity and sets the watermark to its last_event_id. Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Ports DeepInfra's standalone KV-indexer patches from
mm-aware-hash(forked from upstreamb2b786809, 2026-06-02, 2,089 commits behind) onto upstream v1.5.0 (b83b1d930), so the indexer picks up upstream's fixes. One image now serves all three flavors (reality / routing / h24); h24 is the existing--h24runtime flag, not a separate branch.Fixes DEE-722 · Refs DEE-720
Draft, and no live indexer is swapped. Swapping is Thach's call; the rollout proposal is below.
What this branch is
rebase-v1.5.0=v1.5.0+ 3 upstream fixes frommainthat landed after the 09-01 release branch + 10 port commits + 1 test commit. Every port commit names the originals it ports.fix(kv-router): prevent lookup repair…(ai-dynamo#11891),release unowned radix branches…(ai-dynamo#14878),make standalone indexer cleanup safe(ai-dynamo#14652)maincherry-picks with-x. ai-dynamo#14652 conflicted in one selection test only.cd18634d7(image),ef5b033ac,f4b4f1d20; addslibprotobuf-dev, which v1.5.0's sidecar protos needae7403cbe(indexer SUB socket only)959e440f0,90409c7f8(a listener guard;convert_eventandlib/llmuntouched)/kv_recoverwith timeout, gate, retryab351013f,37e1b1a26(arecover_endpointnext to upstream's ZMQreplay_endpoint, which is unchanged)cd18634d7,d14e05d6d(model-name watch),b662110cc(engine_hash → routing group),bb6dcfd5c(DP ranks)b9a2249c4b662110ccd14e05d6d,c9eabbb46; supersedes6d16bb8a95e5f3b8dd73474e51a,247ac26adAudit (all 18 + 2 h24 commits)
6d16bb8a9--ignore-evictions/--merge-shards5e5f3b8ddcd18634d75e5f3b8ddef5b033acab351013f/kv_recoverrecoverylistener.rs:473-506)d14e05d6dc9eabbb46b9a2249c4MooncakeOverlapSummary(services/overlap.rs:18) has none959e440f0convert.rs:432)b662110cc80214ed88lower_tier.rs→Ok/continue)90409c7f8bb6dcfd5c37e1b1a26ae7403cbecommon/zmq.rssets no HWMf4b4f1d20c707ba7ad61764d853image_token_id= image-id mixing (convert.rs:375)73474e51a--h24backend247ac26adUpstream fixes gained: ai-dynamo#10368 (
"CPU"→ HostPinned,protocols.rs:882) and ai-dynamo#12456 (fail closed on unknown media,zmq_wire/mod.rs:199) come with v1.5.0. ai-dynamo#11891, ai-dynamo#14878 and ai-dynamo#14652 are cherry-picked frommain.Behaviour changes vs the images running today
"CPU"KV events go to the HostPinned tier (upstream feat(kv-router): Route vLLM CPU KV events to HostPinned and count lower-tier applies ai-dynamo/dynamo#10368, in v1.5.0). The fork mapped any unknown medium, including vLLM's"CPU", to the device tree. On CPU-offload pods,/querytherefore reported CPU residency asgpu, and CPUBlockRemovedevents removed device entries. deepapi'skv_shards_from_indexer_responsemarks a match SECONDARY whengpu < longest_matched, so routing weights change on CPU-offload models. This is the intended fix.routing_groupfans out over every routing group of the model (as before). A request that namesrouting_groupnow reads only that tree, as upstream does. deepapi sends onlytenant_id, so it is unaffected. Upstream'squery_by_hash_isolates_routing_groups_and_ignores_legacy_tenant_idis renamed and updated for this.--image-token-idis set, and what61764d853restores. Shang-Pin confirmed (a) is the intended default. V4.1-Flash reality/routing currently runc707ba7ad(token-only); swapping them requires deepinfra/backend#5428 to land first.--watch-*in a build withoutkube-discoveryis now an error instead of a warning.--tenant-idis hidden and ignored upstream (use--routing-group). No prod deployment passes it.Tests
cargo test -p dynamo-kv-router --features standalone-indexer,metrics,kube-discovery --lib --test standalone_indexer_http: 1045 passed, 2 ignored; HTTP 7 passed. v1.5.0 baseline: 996.cargo checkof the python bindings (--features kv-indexer-metrics). The one exception: fix(kv-router): prevent lookup repair from restoring removed entries ai-dynamo/dynamo#11891 on its own had one flaky failure that did not reproduce in 4 reruns.cargo clippy … --all-targets -D warnings: only the 3 findings v1.5.0 already has with these features (from_selected_worker_tiers,configure_send_socket,create_bound_pub_socket)."CPU"medium → HostPinned through the listener;/kv_recoverretry-until-success, give-up after 3 attempts, per-attempt timeout;/kv_recoverresponse captured from a prod GLM-5.3 engine deserializes;/query_by_hashcontract (pod_name,longest_matched,gpu);Shadow validation
Image
localhost:30500/dynamo-indexer:kvtest-812b1101cbran as an extra reality-flavor indexer per model: the same args as the livekv-indexer:reality, with no Service and nodi/model_servicelabel, so deepapi never queried it. It ran 2026-09-23 23:17 → 09-24 01:59 UTC (2 h 40 min) and was then deleted.Probes were built from the engines' own ZMQ KV-event stream: every stored prefix traced back to a root, hashed as deepapi does. Each probe went to
/query_by_hashon the live reality indexer and on the shadow back-to-back, half fresh (just stored) and half aged 5–120 min. Probe tool, raw results and shadow manifests:di-sjc-25:/data/home/thach/scratch/dee722-shadow/."CPU"stores and removes (~5.2M of 5.7M stores in the window were non-GPU). On those pods the shadow keeps the CPU chain in HostPinned where the live fork loses it. Example: on pod…-nqezthe live indexer showed 209,024 tokens (allgpu), while the shadow showedgpu209,024 +cpu398,272.kv_recover request failedin any window on any model. All V4.1-Flash ranks (30 pods × 2) recovered at startup through the gate.unknown mediumdrops, partial-prefix skips, block-size exits.Rollout proposal (Thach's call)
c707today.DeepApi:kv_tokens{source="h24"}resumes within 15 min, then the rest.kv_routing_use_indexeris set. Start with a non-offload model (GLM-5.3-Flash), then V4.1-Flash (after step 1), then CPU-offload models (GLM-5.3), where routing weights change by design: CPU matches become SECONDARY. Before and after each swap, comparereality_bestagainstactualonDeepApi:kv_tokensfor 1 h./kv_recoverdumps are applied. The shadows logged every rank recovered within ~1 min and matched live on the first probes after the 10-min warm-up; the brief's ~15 min is the conservative figure.kvtest-80214ed88a/kvtest-c707ba7adb/kvtest-247ac26adc, per the service's current image). Nothing is persisted, so rollback costs the same rebuild.Not done / known gaps
d14e05d6d's CLI init. No deployment passes--enable-logging.evictions.rsandh24.rsstill read cgroup memory files and the clock directly (ported as-is, not redesigned).lib/llmworker-query recovery, which the standalone indexer never runs.extra_keys) stores are not probed.https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa