Skip to content

standalone-indexer: port DeepInfra patches onto upstream v1.5.0 (DEE-722) - #49

Draft
Thachnh wants to merge 14 commits into
upstream-v1.5.0from
rebase-v1.5.0
Draft

Thachnh wants to merge 14 commits into
upstream-v1.5.0from
rebase-v1.5.0

Conversation

@Thachnh

@Thachnh Thachnh commented Sep 24, 2026 •

Copy link
Copy Markdown

Ports DeepInfra's standalone KV-indexer patches from mm-aware-hash (forked from upstream b2b786809, 2026-06-02, 2,089 commits behind) onto upstream v1.5.0 (b83b1d930), so the indexer picks up upstream's fixes. One image now serves all three flavors (reality / routing / h24); h24 is the existing --h24 runtime flag, not a separate branch.

Fixes DEE-722 · Refs DEE-720

Draft, and no live indexer is swapped. Swapping is Thach's call; the rollout proposal is below.

What this branch is

rebase-v1.5.0 = v1.5.0 + 3 upstream fixes from main that landed after the 09-01 release branch + 10 port commits + 1 test commit. Every port commit names the originals it ports.

commit ports
fix(kv-router): prevent lookup repair… (ai-dynamo#11891), release unowned radix branches… (ai-dynamo#14878), make standalone indexer cleanup safe (ai-dynamo#14652) upstream main cherry-picks with -x. ai-dynamo#14652 conflicted in one selection test only.
indexer image: build from a git archive cd18634d7 (image), ef5b033ac, f4b4f1d20; adds libprotobuf-dev, which v1.5.0's sidecar protos need
unbounded ZMQ receive queue ae7403cbe (indexer SUB socket only)
exit on block-size mismatch, skip partial-prefix stores 959e440f0, 90409c7f8 (a listener guard; convert_event and lib/llm untouched)
HTTP /kv_recover with timeout, gate, retry ab351013f, 37e1b1a26 (a recover_endpoint next to upstream's ZMQ replay_endpoint, which is unchanged)
discover engine pods from Kubernetes cd18634d7, d14e05d6d (model-name watch), b662110cc (engine_hash → routing group), bb6dcfd5c (DP ranks)
pod_name through /workers and /query b9a2249c4
/query from every routing group query half of b662110cc
--keep-evictions d14e05d6d, c9eabbb46; supersedes 6d16bb8a9
--enable-logging kv_audit 5e5f3b8dd
--h24 73474e51a, 247ac26ad

Audit (all 18 + 2 h24 commits)

# ours what v1.5.0 status result
1 6d16bb8a9 --ignore-evictions / --merge-shards superseded by keep-evictions; merge-shards removed in 5e5f3b8dd dropped
2 cd18634d7 k8s pod watcher, CLI, image files no discovery upstream ported
3 5e5f3b8dd kv_audit logging upstream access log (ai-dynamo#10700) has no hashes/events ported
4 ef5b033ac release build ours ported
5 ab351013f HTTP /kv_recover recovery upstream is ZMQ replay only (listener.rs:473-506) ported
6 d14e05d6d model-name watch + keep-evictions not upstream ported
7 c9eabbb46 keep-evictions status line not upstream ported
8 b9a2249c4 pod_name in /workers, /query MooncakeOverlapSummary (services/overlap.rs:18) has none ported
9 959e440f0 fatal block-size mismatch v1.5.0 still drops silently (convert.rs:432) ported (listener guard)
10 b662110cc per-engine_hash trees + fan-out not upstream ported
11 80214ed88 lower-tier removes skip unknown hashes covered by ai-dynamo#14652 (lower_tier.rs → Ok/continue) dropped
12 90409c7f8 skip partial-prefix stores not upstream ported
13 bb6dcfd5c one listener per DP rank not upstream ported
14 37e1b1a26 recover timeout/gate/retry v1.5.0 has no HTTP recovery ported
15 ae7403cbe RCVHWM=0 common/zmq.rs sets no HWM ported
16 f4b4f1d20 touch sources in image ours ported
17 c707ba7ad token-only image hashing reverted by #18 dropped
18 61764d853 revert of #17 net zero; v1.5.0 default without image_token_id = image-id mixing (convert.rs:375) nothing to port
h1 73474e51a --h24 backend not upstream ported
h2 247ac26ad h24 one tree per model not upstream ported

Upstream fixes gained: ai-dynamo#10368 ("CPU" → HostPinned, protocols.rs:882) and ai-dynamo#12456 (fail closed on unknown media, zmq_wire/mod.rs:199) come with v1.5.0. ai-dynamo#11891, ai-dynamo#14878 and ai-dynamo#14652 are cherry-picked from main.

Behaviour changes vs the images running today

  1. "CPU" KV events go to the HostPinned tier (upstream feat(kv-router): Route vLLM CPU KV events to HostPinned and count lower-tier applies ai-dynamo/dynamo#10368, in v1.5.0). The fork mapped any unknown medium, including vLLM's "CPU", to the device tree. On CPU-offload pods, /query therefore reported CPU residency as gpu, and CPU BlockRemoved events removed device entries. deepapi's kv_shards_from_indexer_response marks a match SECONDARY when gpu < longest_matched, so routing weights change on CPU-offload models. This is the intended fix.
  2. Unknown media are dropped (fix(kv-router): fail closed on unrecognized KV event media ai-dynamo/dynamo#12456) instead of being filed as device.
  3. A /query without a routing_group fans out over every routing group of the model (as before). A request that names routing_group now reads only that tree, as upstream does. deepapi sends only tenant_id, so it is unaffected. Upstream's query_by_hash_isolates_routing_groups_and_ignores_legacy_tenant_id is renamed and updated for this.
  4. Image blocks keep image-id mixing: v1.5.0's default when no --image-token-id is set, and what 61764d853 restores. Shang-Pin confirmed (a) is the intended default. V4.1-Flash reality/routing currently run c707ba7ad (token-only); swapping them requires deepinfra/backend#5428 to land first.
  5. Configuring --watch-* in a build without kube-discovery is now an error instead of a warning.
  6. --tenant-id is hidden and ignored upstream (use --routing-group). No prod deployment passes it.

Tests

  • cargo test -p dynamo-kv-router --features standalone-indexer,metrics,kube-discovery --lib --test standalone_indexer_http: 1045 passed, 2 ignored; HTTP 7 passed. v1.5.0 baseline: 996.
  • Every commit on the branch passes that suite plus cargo check of the python bindings (--features kv-indexer-metrics). The one exception: fix(kv-router): prevent lookup repair from restoring removed entries ai-dynamo/dynamo#11891 on its own had one flaky failure that did not reproduce in 4 reruns.
  • cargo clippy … --all-targets -D warnings: only the 3 findings v1.5.0 already has with these features (from_selected_worker_tiers, configure_send_socket, create_bound_pub_socket).
  • New tests include:
    • "CPU" medium → HostPinned through the listener;
    • /kv_recover retry-until-success, give-up after 3 attempts, per-attempt timeout;
    • a /kv_recover response captured from a prod GLM-5.3 engine deserializes;
    • a TreeDump gap fill end to end;
    • deepapi's golden block hashes (plain + LoRA) and the /query_by_hash contract (pod_name, longest_matched, gpu);
    • routing-group fan-out, keep-evictions parking, h24 one-tree-per-model.

Shadow validation

Image localhost:30500/dynamo-indexer:kvtest-812b1101cb ran as an extra reality-flavor indexer per model: the same args as the live kv-indexer:reality, with no Service and no di/model_service label, so deepapi never queried it. It ran 2026-09-23 23:17 → 09-24 01:59 UTC (2 h 40 min) and was then deleted.

Probes were built from the engines' own ZMQ KV-event stream: every stored prefix traced back to a root, hashed as deepapi does. Each probe went to /query_by_hash on the live reality indexer and on the shadow back-to-back, half fresh (just stored) and half aged 5–120 min. Probe tool, raw results and shadow manifests: di-sjc-25:/data/home/thach/scratch/dee722-shadow/.

model probes equal best match shadow more / less matched tokens shadow ÷ live best pod agrees
zai-org/GLM-5.3 (block 64, CPU offload) 6948 6785 138 / 25 1.018 5438 exact, ties dominate the rest
deepseek-ai/DeepSeek-V4.1-Flash (DP2, block 128) 7392 7378 13 / 1 1.000 7317
zai-org/GLM-5.3-Flash (hybrid, block 2432) 7947 7930 14 / 3 1.000 7280
  • CPU tier: 2 of the 7 GLM-5.3 pods publish "CPU" stores and removes (~5.2M of 5.7M stores in the window were non-GPU). On those pods the shadow keeps the CPU chain in HostPinned where the live fork loses it. Example: on pod …-nqez the live indexer showed 209,024 tokens (all gpu), while the shadow showed gpu 209,024 + cpu 398,272.
  • Recovery: 0 kv_recover request failed in any window on any model. All V4.1-Flash ranks (30 pods × 2) recovered at startup through the gate.
  • ParentBlockNotFound, shadow vs live over the last 2 h: GLM-5.3 0 vs 1592 (the shadow's 2667 all came in the first minute, while dumps were applying); V4.1-Flash 1387 vs 3807; GLM-5.3-Flash 460 vs 293. No flood.
  • Memory (cgroup, first → last): GLM-5.3 shadow 342 → 421 MiB (live 686 → 1128); V4.1-Flash shadow 3334 → 3433 MiB (live 5521 → 5363); GLM-5.3-Flash 45 → 48 MiB (live 219 → 220). 0 restarts.
  • Not seen at all: unknown medium drops, partial-prefix skips, block-size exits.

Rollout proposal (Thach's call)

  1. Merge-blockers first: land deepinfra/backend#5428 (image-id-mixed probe) before any V4.1-Flash swap. Its reality/routing indexers run token-only c707 today.
  2. h24 flavor first (measurement only; deepapi reads it for the ladder, never for routing): one model, e.g. zai-org/GLM-5.3. Verify DeepApi:kv_tokens{source="h24"} resumes within 15 min, then the rest.
  3. routing flavor (keep-evictions; measurement only), model by model.
  4. reality flavor last. This is what deepapi routes with when kv_routing_use_indexer is set. Start with a non-offload model (GLM-5.3-Flash), then V4.1-Flash (after step 1), then CPU-offload models (GLM-5.3), where routing weights change by design: CPU matches become SECONDARY. Before and after each swap, compare reality_best against actual on DeepApi:kv_tokens for 1 h.
  5. Mechanics: an image bump in each service's manifest and a redeploy. A redeploy wipes that indexer's tree until the /kv_recover dumps are applied. The shadows logged every rank recovered within ~1 min and matched live on the first probes after the 10-min warm-up; the brief's ~15 min is the conservative figure.
  6. Rollback: redeploy the previous tag (kvtest-80214ed88a / kvtest-c707ba7adb / kvtest-247ac26adc, per the service's current image). Nothing is persisted, so rollback costs the same rebuild.

Not done / known gaps

  • Not ported: the non-blocking stdout log writer from d14e05d6d's CLI init. No deployment passes --enable-logging.
  • evictions.rs and h24.rs still read cgroup memory files and the clock directly (ported as-is, not redesigned).
  • upstream fix(kv-router): restore server-selected buffer-first recovery ai-dynamo/dynamo#14627 is not cherry-picked: it only changes lib/llm worker-query recovery, which the standalone indexer never runs.
  • Only the reality flavor was shadowed. keep-evictions and h24 are covered by unit tests only.
  • The shadow probes are built from prefixes stored while the probe ran (fresh, and aged 5–120 min). Older prefixes and image-bearing (extra_keys) stores are not probed.
  • Size: ~3.6k added lines of ours, most of it ported modules with their unit tests (pod_watcher, evictions, h24, kv_recover). It is split one commit per original patch for review.

https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa

jthomson04 and others added 14 commits September 23, 2026 22:12
…i-dynamo#11891)

Signed-off-by: jthomson04 <64760228+jthomson04@users.noreply.github.com>
(cherry picked from commit 6b78d18)
…amo#14878)

Signed-off-by: jthomson04 <jwillthomson19@gmail.com>
(cherry picked from commit 8f95101)
Signed-off-by: PeaBrane <yanrpei@gmail.com>
(cherry picked from commit 6c66d4b)
Carries forward DeepInfra's indexer image recipe: a two-stage Ubuntu
24.04 build that compiles the dynamo._core wheel with kv-indexer-metrics
and installs only that wheel into a slim runtime.

Ports cd18634 (image + build script), ef5b033 (release build) and
f4b4f1d (touch archived sources so a sibling branch's newer artifact in
the shared cargo cache is never reused). Header comment now lists the
three service flavors the one image serves.

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
A listener blocked in gap recovery can let its SUB pipe reach the default
RCVHWM of 1000; with heartbeats on, libzmq 4.3.4 then aborts on
`Assertion failed: _input_stopped` (zeromq/libzmq#3596, ai-dynamo#3937). Seen in
prod as an h24 indexer crash loop while 66 startup TreeDumps applied.

Only the indexer's connect_sub_socket changes; the shared receive config
used by replica sync, selection and the slot tracker keeps its default.

Ports ae7403c.

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
…x stores

The shared normalizer drops every block whose size differs from the
configured one, so an indexer started with the wrong --block-size stays
empty while its stream looks alive. The standalone listener now checks
each preprocessed BlockStored first:
- a store of blocks larger than configured is a misconfiguration: log and
  exit so k8s surfaces a CrashLoopBackOff;
- a store smaller than configured is a vLLM partial-prefix entry
  (prefix_match_unit < cache block, e.g. 32/64/96 of 128 on V4.1): drop it,
  logging the first three.

Ports 959e440 and 90409c7. Unlike the originals this lives in the
standalone listener (new block_size.rs) and leaves convert_event's
signature and the lib/llm publisher untouched; the publisher half of
959e440 is Dynamo's own worker-side path, which we don't run.

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
…gate and retry

Our engines serve their local indexer at GET /kv_recover (HTTP) rather
than on a ZMQ ROUTER replay socket. A listener registered with a
recover_endpoint now recovers a sequence gap by fetching a
WorkerKvQueryResponse for [start, end) and applying it:
- Events: apply all under the consumer-assigned (worker, dp_rank),
  resume at last_event_id;
- TreeDump (whole-rank scope): remove_worker_dp_rank, apply all, resume at
  last_event_id; a domain-scoped dump applies without the reset;
- TooNew / InvalidRange / Error / TreeDumpFailed: no-op.
Without a recover_endpoint the upstream ZMQ replay path is unchanged.

Downloads share one reqwest client and a process-wide gate
(--recover-concurrency, default 8), with a total timeout of
--recover-timeout-secs (default 120) and 3 attempts with 2/4 s backoff
before the gap is given up ("kv_recover request failed; giving up").
v1.5.0 has no HTTP recovery at all (its only HTTP client is peer /dump
recovery, hard-coded to 10 s).

The settings are injected through WorkerRegistry::with_kv_recover rather
than process globals, and /register takes recover_endpoint as a new field
next to replay_endpoint instead of aliasing it. register() keeps
upstream's signature; register_with_extras() carries our fields.

Tests: a /kv_recover response captured from a prod GLM-5.3 engine
deserializes and plans under the consumer identity; retry until success;
give up after the last attempt; per-attempt timeout.

Ports ab35101 and 37e1b1a.

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
With --watch-namespace plus --watch-model-name (or a raw --watch-label),
the indexer watches the model's engine pods and keeps the registry in
sync: a Ready pod is registered, a deleted / unready pod deregistered.
kube::runtime's reflector re-lists on 410 Gone, and every change
reconciles the desired Ready set against what is registered, so a missed
delete is corrected on the next re-list.

- The selector is di/model_name=<sanitized model name>, the label the
  backend stamps on every engine pod across engine_hash variants.
- Each pod registers under routing_group = its engine_hash label (the
  configured group if unlabeled), so hash-incompatible engine generations
  never share a tree during a migration; a changed label re-registers the
  pod under its new group.
- --watch-dp-size N registers one listener per data-parallel rank:
  dp_rank r on tcp://<ip>:<zmq_port + r>, recovering from
  http://<ip>:<recover_port + r>. A pod whose registration fails part-way
  is deregistered before the retry.
- Configuring discovery in a build without the kube-discovery feature is
  now an error instead of a warning.

Discovery config and the label helpers (discovery.rs) are always compiled;
the kube client lives only in the feature-gated pod_watcher.rs. The
python bindings' kv-indexer feature enables kube-discovery.

Ports cd18634 (watcher), the model-name watch of d14e05d, the
engine_hash grouping of b662110, and bb6dcfd (DP ranks).

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
Workers are keyed by instance_id (xxh3 of the pod name), which is one
way. deepapi maps /query_by_hash matches back to shards by pod name
(kv_shards_from_indexer_response), so keep the raw pod name on the worker
record when the pod watcher or a /register caller provides it and surface
it in /workers entries and in each /query `instances` entry.

v1.5.0's shared MooncakeOverlapSummary has no pod_name; the indexer
server wraps it in a flattened InstanceMatch rather than changing the
shared type. Purely additive to the response.

Adds deepapi_contract_tests.rs: the /query_by_hash request/response
shape deepapi depends on.

Ports b9a2249.

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
The pod watcher gives each engine_hash regime its own tree (routing
group). A /query or /query_by_hash that names no routing_group now fans
out over every tree of model_name and merges the per-tree responses:
worker sets are disjoint, so scores and instances union; frequencies sum
element-wise. A probe hashed in one regime scores 0 against the others,
so routing converges on pods that can reuse the cache during image
migrations. A request that names a routing_group still reads only that
tree, as upstream does; the legacy tenant_id stays ignored. deepapi
sends only tenant_id "default", so it gets the fan-out.

A failing tree is skipped while any tree answers (500 only if all fail);
404 only when the model has no tree.

The fan-out lives in model_query.rs and replaces upstream's single-tree
run_tiered_query. Upstream's
query_by_hash_isolates_routing_groups_and_ignores_legacy_tenant_id is
renamed and its no-routing_group case now expects the fan-out. New tests:
/query_by_hash across two groups excluding another model, /query by
tokens, and deepapi's golden block hashes (plain and LoRA-seeded) from
deepinfra/utils.kv_block_hashes.

Ports the query half of b662110.

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
…mory pressure

The kv-indexer:routing flavor measures "infinite memory, per-shard" hit
rates: engine Removed events are parked in a per-listener buffer instead
of applied (Cleared is dropped), and a store of a parked block cancels
its eviction. Every 15 s, once memory use crosses
--evict-memory-threshold (default 0.75) of the cgroup limit, evictions
older than --evict-retention-secs (default 1800) are replayed into the
tree as ordinary Removed events, so the pod degrades toward reality
instead of OOMing. A per-minute status line logs memory fraction,
pending and queued entries, and listener count.

A replayed eviction is applied only if it is still the block's latest;
a /kv_recover TreeDump that resets the rank clears the rank's buffer.
The flag is a RegistryOptions field (no process globals): only listeners
of a keep-evictions registry carry a buffer.

Ports the keep-evictions half of d14e05d and c9eabbb; supersedes
6d16bb8 (--ignore-evictions / --merge-shards).

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
…s and events

Off by default. When on, the indexer logs one kv_audit line per /query
and /query_by_hash (model, routing groups queried, status, probe block
hashes, full JSON response) and one per engine event it ingests (STORE /
EVICT / CLEAR with worker, dp_rank, tier, event id, seq, and the token
and sequence block hashes), tagged source=live|replay|recover and written
before the --keep-evictions filter so it shows what the engine published.
Isolate with RUST_LOG=kv_audit=info. Upstream's --access-log (ai-dynamo#10700)
records requests but not hashes, responses or events.

Ports 5e5f3b8. Not ported: the non-blocking stdout writer from
d14e05d's CLI init; no prod deployment passes --enable-logging.

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
A new Indexer::H24 backend answers "how many prefix blocks of this query
were stored by ANY worker within the retention horizon", ignoring
evictions, clears and worker identity. Two flat maps replace the radix
tree: engine sequence hash -> chained prefix hash (so stores naming their
parent by engine hash keep chaining) and prefix hash -> last_stored /
last_touched. Matches are reported under a synthetic worker 0, so the
/query_by_hash response shape is unchanged.

- Matched-block ages feed dynamo_kvindexer_h24_age_seconds, whose bucket
  CDF over queried blocks is the hit-rate-vs-retention curve; queried
  blocks are counted once per request, not per tree.
- An expiry sweep every 10 min drops entries older than
  --h24-horizon-secs (default 172800), and tightens to half the horizon
  above 92% of the cgroup memory limit.
- Every routing group of a model collapses into one h24 tree, so a block
  matched in several engine_hash regimes is counted once (the 120%
  hit-rate incident).
- /dump returns nothing in h24 mode; --h24 and --keep-evictions are
  mutually exclusive.

A runtime flag of the one image rather than a separate -h24 branch:
every live h24 deployment already passes --h24.

Ports 73474e5 and 247ac26.

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
Covers, through ListenerLoop with wire-format msgpack batches:
- "CPU" medium stores land on the HostPinned tier, not the device tree
  (upstream ai-dynamo#10368; the fork this replaces filed them as device);
- sub-block partial-prefix stores are skipped without exiting;
- --keep-evictions parks an eviction and a re-store cancels it;
- a sequence gap is filled from /kv_recover: a TreeDump replaces the
  rank's state under the consumer identity and sets the watermark to its
  last_event_id.

Claude-Session: https://claude.ai/code/session_01XKZ49SpoymBqDqgGRJYMNa
@Thachnh
Thachnh deployed to external_collaborator September 24, 2026 02:01 — with GitHub Actions Active
@github-actions github-actions Bot added documentation Improvements or additions to documentation container labels Sep 24, 2026

This branch was successfully deployed

1 active deployment
external_collaborator — 812b1101 Deployed Sep 24, 2026 by Thachnh via ok-to-test #55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

container documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants