Repository navigation
perf(#1024): lazy request-cache seen-set — 206 to 64 kB/idx idle RSS - #1025
Merged
Merged
Conversation
The #1024 idle-RSS audit reconstructed the live budget of one open index at ~205 kB resident (4-vCPU idle-budget arm, flat in N; 183-194 kB/idx on the 32-vCPU arm at N>=300). The single largest line item is the request-cache seen-set: RequestCacheSeen::with_capacity(65_536) ran at BOTH Index construction sites (the create and the open struct literals), pre-writing a 131,072-bucket HashSet table plus a 512 KiB VecDeque ring - measured 135 kB + ~1 kB resident per set under the server's jemalloc config (background_thread:true, dirty_decay_ms:1000), i.e. ~136 of ~205 kB/idx, paid by every index at boot or create before it has ever served a request_cache=true search. The fixture node's 14 system indices pay it too. Change (lazy allocation, no restructuring): RequestCacheSeen::new(cap) starts both collections empty (HashSet::default() / VecDeque::default() - zero bytes, zero written pages); record() grows them exactly as an unsized HashSet/VecDeque grows, and the FIFO bound (evict oldest at cap) is enforced exactly as before. First sight = miss, repeat = hit, unchanged; the ES wire protocol is untouched. A busy index converges to the same bounded footprint as before (65,536 entries, ~1.5 MiB logical); an idle or never-request_cache-queried index holds nothing. Design adapted (not copied) from quickwit's pay-per-use reader stack: quickwit-search/src/leaf.rs:237-283 open_index_with_caches (Apache-2.0) opens the whole per-split Index inside the leaf-search future and drops it with the query, keeping idle memory at zero per split - the same zero-until-used shape, applied to per-index bookkeeping. Fail-before tests (crates/xerj-engine/src/index.rs, request_cache_seen_idle_tests; ci-test profile, this tree): - created_index_allocates_no_request_cache_seen_set FAILED before the change (HashSet capacity 131,072 at create), ok after - reopened_index_allocates_no_request_cache_seen_set FAILED before (same on the boot/open path; also asserts the reopened index still serves its document, so laziness provably loses nothing), ok after - track_request_cache_dedupes_and_evicts_when_lazily_grown pins the hit/miss contract and FIFO eviction at the cap through the public path (65,537 distinct hashes, oldest evicted and re-missed), ok before and after Gates: xerj-engine --lib 730 passed / 0 failed; xerj-api suite all green; ES-YAML conformance 1376 passed / 0 failed / 3 skipped (private port, throwaway dir); rustfmt clean; clippy -D warnings clean. Measured (idle-budget fixture, same host, back-to-back A/B on the before-binary preserved from ed510a6, allocator pins thp:never,dirty_decay_ms:0,muzzy_decay_ms:0): - 4-vCPU taskset 0-3, N=450: 203.2 -> 68.0 kB/idx (-66%; RssAnon 108,008 -> 45,028 kB). Idle CPU identical (0.1583% both arms), wakeups 25.3 -> 25.1/s, boot-to-green 6 ms both, 0 WAL-replay lines both. - 32-vCPU, N=300: 186.4 -> 51.7 kB/idx (-72%; RssAnon 103,488 -> 60,736 kB), boot 8 -> 7 ms, 0 WAL-replay lines both. - The 32-vCPU after-run tripped the idle-CPU gate line (0.766% > 0.5%) with loadavg 28.99 at window start - foreign tenants on the shared unpinned box (the before-run read 0.475% at loadavg 5.39). A second attempt also landed on a loaded window (0.675%, loadavg 21.68; RSS 54.3 kB/idx, consistent). The controlled comparison is the taskset'd 4-vCPU pair, which read byte-identical idle CPU (0.1583% both arms) and wakeups 25.3 -> 25.1/s: this commit adds no timers or polling and strictly removes boot-time ctrl-page writes, so the unpinned-arm flap is box noise, not the change. Wakeups on the 32-vCPU arm stayed 25.6 -> 28.0/s. Every other gate line passed on every run. Issue #874's product line (204.8 kB/idx) is now met on both host shapes with 3.0-4.0x margin. Follow-ups (next tranches; the issue stays open): the remaining ~50-70 kB/idx is dominated by ~24 per-index DashMap 16-shard arrays (~2.1 kB each, eagerly written at construction), WAL/memtable shard structs, and 3-4 parsed JSON copies of settings/schema/mapping; the cold-path endgame is quickwit-style idle demotion (quickwit-ingest/src/ingest_v2/ idle.rs:25-67 CloseIdleShardsTask + models.rs:182-197 is_idle, Apache-2.0 - one last_write_instant timestamp and a weak-state periodic closer) driving the existing #463 release path by idle age for cleanly-flushed indexes. Refs #1024.
Appends the after-numbers to benchmarks/idle-budget/README.md's RSS calibration section (history preserved; nothing rewritten) and commits the result files they cite, verifying commit 6526e8ea (lazy request-cache seen-set) on this host with the idle-budget fixture, ci-test binaries and allocator pins on. 4-vCPU (taskset -c 0-3), per-index idle RSS (VmRSS minus empty node): N=150 before 204.9 kB -> after 63.2 kB (-69%) N=450 before 203.7 kB -> after 63.3 kB (-69%) N-slope (450-arm minus 150-arm)/300, VmRSS: 206.5 -> 63.7 kB/idx (-69%) (RssAnon slope: 205.7 -> 68.8). The subtraction and the slope now agree: the curve is flat at a genuinely O(N)-constant ~64 kB/idx, and issue #874's 0.2 MB/index product line is met with >3x margin. Idle CPU at N=450 0.225% -> 0.192%, wakeups 25.2 -> 21.4/s, boot 6 -> 6 ms, WAL replay 0 both. 32-vCPU (unpinned), N=450 — before arm added for this verification (result-p1024-before-32cpu-n450.json; every earlier 32-vCPU reading was N=300, so the shape had no N=450 baseline): per-index 199.0 kB -> 66.3 / 62.6 kB (two after runs; -67%) idle CPU 0.767% (loadavg 19) -> 0.650 / 0.633% (loadavg 3-4) wakeups 22.6 -> 25.9 / 22.2 per s The >0.5% idle-CPU gate line trips on the BEFORE binary too on this arm (0.767%): pre-existing for the unpinned many-core shape at N=450 on this shared box, not a regression from 6526e8ea - every matched pair reads lower after. The controlled 4-vCPU pair is the comparison; wakeups (the timer-churn proxy) are flat on every arm. Gates re-run for this verification: cargo fmt --check clean; clippy --profile ci-test -p xerj-engine --all-targets -- -D warnings clean; cargo test --profile ci-test -p xerj-engine 1221 passed / 0 failed / 0 ignored; ES-YAML conformance 1376 passed / 0 failed / 3 skipped (private port 9630, throwaway data dir). Refs #1024.
xerj-org
added a commit
that referenced
this pull request
Sep 26, 2026
Twelve PRs have merged since the rc.77 tag; the [Unreleased] section carried only the #874 gauges entry. The release-notes gate (.github/scripts/release-notes-gate.sh, issue #474) fails an rc.78 release PR whose notes do not cite every PR in the previous-tag..head window, so each of these would have surfaced as a gate failure at cut time instead of a review comment now. Entries added (with PR links so the gate's coverage check resolves): - Fixed: #1015 flush-drain freeze (PR #1018), #950 id-position maps + streamed reassembly (PR #1017), #1019 _delete_by_query paging (PR #1021), #1022 _update_by_query paging (PR #1023); the existing #874 entry now cites its PR (#1020). - Added: POST /{index}/_cache/clear (PR #1009), the systemone email-labelling benchmark answering discussion #1012 (PR #1026) with its measured numbers (1.000 templated / 0.625 at 0.902 confidence on the hard tier). - Performance: request-cache seen-set lazy allocation, idle 206 -> 64 kB/idx (PRs #1025, #1034). - Documentation: README Jev section (PR #1010), the /_decide field report (PR #1027), llms.txt status catch-up (PR #1033, closing #1028). Numbers are quoted only from the PR bodies' own verified runs.
Amakurai
pushed a commit
to Amakurai/xerj
that referenced
this pull request
Sep 27, 2026
…, rc.78 queue The 2026-09-21 review of this file was stale within hours: xerj-org#950 closed 19:33, xerj-org#941 19:38, xerj-org#1015 20:49 (all 2026-09-21), xerj-org#874 2026-09-22, and the stage-1 gating trio xerj-org#937/xerj-org#938/xerj-org#939 closed the same evening — every one recorded here as open or "under way". This pass re-checks every status claim against live tracker state and rolls the file forward: - Next release: the two "in flight" items (xerj-org#1015, xerj-org#950) are closed with fixes on main riding to rc.78 — the section now records what actually landed since the rc.77 tag (xerj-org#1009, xerj-org#1017, xerj-org#1018, xerj-org#1020, xerj-org#1021, xerj-org#1023, xerj-org#1025, xerj-org#1026, xerj-org#1033, xerj-org#1034) with PR links, including xerj-org#874's budget met at ~3x margin (206 -> 64 kB per idle index) and discussion xerj-org#1012's email-labelling measurement (1.000 templated tier / 0.625 at 0.902 confidence on the hard tier — wrong-and-confident). - Open defects: now xerj-org#1031 and xerj-org#1032 only; xerj-org#1015/xerj-org#950 removed (closed); trackers sentence corrected (xerj-org#941, xerj-org#874 closed — only xerj-org#298 remains by design). - GA gate: "Close the CHANGELOG gap" marked closed 2026-09-26 (the rc.19-rc.70 backfill, PR xerj-org#1035); the record-complete note replaces the gap warning under "Shipping today". - Zero-token: stage-1 gating trio recorded as fixed (xerj-org#991, xerj-org#995, xerj-org#979) with the old measured costs kept as the before-state; the xerj-org#940 hash-seed spread annotated as a pre-fix measurement; stage-2 object storage marked done (xerj-org#965 wired in rc.77 — the old bullet contradicted this file's own "Shipping today" section); mail ingest "unreleased" corrected to "shipped in rc.75"; three stale "In flight:" labels on landed stage-1 items relabelled "History:". - Mail-ingest memory line: xerj-org#948's runaway is fixed (xerj-org#1002) — the line now carries the fixed numbers and points at xerj-org#1032 for the residual. Review line updated to 2026-09-26 (machine-checked format), scoped honestly as a desk review. The milestones line is true again: rc.78 milestone created, all four open issues triaged (xerj-org#1030/xerj-org#1031/xerj-org#1032 -> rc.78, xerj-org#298 -> GA).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Tranche 1 of the #1024 idle-RSS work: per-index idle memory on 4-vCPU hosts
was sitting at the #874 product line (205-206 kB/index), with an unexplained
N-curve knee on many-core hosts.
Root cause
RequestCacheSeen::with_capacity(65_536)ran eagerly at bothIndexconstruction sites (
Index::openand the create path), so every index —including one that has never served a
request_cache=truesearch — paid~136 kB resident at boot: the
HashSetctrl pages of a 131,072-bucket tableplus the
VecDequering, all written at construction. That is ~2/3 of thewhole ~205 kB/index reconstructed idle budget (remainder: ~50 kB of per-index
DashMapshard arrays, ~15-25 kB WAL writers/memtable structs/settingscopies).
Design adapted from quickwit's lazily-built per-index caches
(
quickwit-search/src/leaf.rs:237-283, Apache-2.0; no code copied).Change
RequestCacheSeen::with_capacity→RequestCacheSeen::new(cap)at bothconstruction sites: both collections start
HashSet::default()/VecDeque::default()(zero bytes, zero written pages) and grow on firstrecord(). FIFO bound, hit/miss semantics, and the steady-state footprint ofa busy index are unchanged; wire protocol untouched.
Fail-before proofs (reverted → red, fixed → green)
created_index_allocates_no_request_cache_seen_set(capacity 131,072 → 0)reopened_index_allocates_no_request_cache_seen_set(boot/open path; alsoasserts a reopened index still serves its doc)
track_request_cache_dedupes_and_evicts_when_lazily_grown(contract + FIFOeviction at cap via 65,537 distinct hashes) — held before and after
Before/after (idle-budget fixture, ci-test, allocator pins on, same host)
4-vCPU (
taskset -c 0-3), per-index idle RSS (VmRSS minus empty node):32-vCPU, N=450: 199.0 → 66.3 / 62.6 kB (two runs). Subtraction and slope now
agree: the curve is flat at a genuinely O(N)-constant ~64 kB/idx — the #874
product line (204.8) is met with >3x margin. The calibration section of
benchmarks/idle-budget/README.mdis appended (history preserved) with theresult files committed under
results/.The N-curve knee, explained as measured
On many-core hosts the marginal slope steps from ~145 kB/idx (0→150) to a
~215-225 plateau (150→450); at 4 vCPU it is flat ~206 throughout. The
reconstructed budget (136 eager seen-set + ~50 DashMaps + ~15-25 misc ≈
205-210) matches both readings — the knee is where the eager allocations stop
being absorbed, not a leak. With the seen-set lazy, the after-curve is flat
~64 on both shapes. (Pre-existing, unchanged: the 0.5 % idle-CPU gate line
trips on 32-vCPU shared-box arms on the BEFORE binary too — every matched
pair reads lower after.)
Gates (rebased on current main)
xerj-engine --lib: 731 passed / 0 failedcargo fmt --check,clippy -D warnings(engine, all targets): cleanNot done here (tranche 2, issue stays open)
The remaining ~64 kB/idx: per-index
DashMapshard-array laziness and aquickwit-style idle-demotion loop (
quickwit-ingest/src/ingest_v2/idle.rs:25-67,models.rs:182-197) driving the existing #463 release path.Refs #1024.