Skip to content

perf(#1024): lazy request-cache seen-set — 206 to 64 kB/idx idle RSS - #1025

Merged
xerj-org merged 2 commits into
mainfrom
fix/issue-1024-idle-rss-per-index
Sep 22, 2026
Merged

xerj-org merged 2 commits into
mainfrom
fix/issue-1024-idle-rss-per-index

Conversation

@xerj-org

Copy link
Copy Markdown
Owner

What

Tranche 1 of the #1024 idle-RSS work: per-index idle memory on 4-vCPU hosts
was sitting at the #874 product line (205-206 kB/index), with an unexplained
N-curve knee on many-core hosts.

Root cause

RequestCacheSeen::with_capacity(65_536) ran eagerly at both Index
construction sites (Index::open and the create path), so every index —
including one that has never served a request_cache=true search — paid
~136 kB resident at boot: the HashSet ctrl pages of a 131,072-bucket table
plus the VecDeque ring, all written at construction. That is ~2/3 of the
whole ~205 kB/index reconstructed idle budget (remainder: ~50 kB of per-index
DashMap shard arrays, ~15-25 kB WAL writers/memtable structs/settings
copies).

Design adapted from quickwit's lazily-built per-index caches
(quickwit-search/src/leaf.rs:237-283, Apache-2.0; no code copied).

Change

RequestCacheSeen::with_capacity → RequestCacheSeen::new(cap) at both
construction sites: both collections start HashSet::default() /
VecDeque::default() (zero bytes, zero written pages) and grow on first
record(). FIFO bound, hit/miss semantics, and the steady-state footprint of
a busy index are unchanged; wire protocol untouched.

Fail-before proofs (reverted → red, fixed → green)

  • created_index_allocates_no_request_cache_seen_set (capacity 131,072 → 0)
  • reopened_index_allocates_no_request_cache_seen_set (boot/open path; also
    asserts a reopened index still serves its doc)
  • track_request_cache_dedupes_and_evicts_when_lazily_grown (contract + FIFO
    eviction at cap via 65,537 distinct hashes) — held before and after

Before/after (idle-budget fixture, ci-test, allocator pins on, same host)

4-vCPU (taskset -c 0-3), per-index idle RSS (VmRSS minus empty node):

before after Δ
N=150 204.9 kB 63.2 kB −69 %
N=450 203.7 kB 63.3 kB −69 %
N-slope VmRSS (450−150)/300 206.5 kB/idx 63.7 kB/idx −69 %

32-vCPU, N=450: 199.0 → 66.3 / 62.6 kB (two runs). Subtraction and slope now
agree: the curve is flat at a genuinely O(N)-constant ~64 kB/idx — the #874
product line (204.8) is met with >3x margin. The calibration section of
benchmarks/idle-budget/README.md is appended (history preserved) with the
result files committed under results/.

The N-curve knee, explained as measured

On many-core hosts the marginal slope steps from ~145 kB/idx (0→150) to a
~215-225 plateau (150→450); at 4 vCPU it is flat ~206 throughout. The
reconstructed budget (136 eager seen-set + ~50 DashMaps + ~15-25 misc ≈
205-210) matches both readings — the knee is where the eager allocations stop
being absorbed, not a leak. With the seen-set lazy, the after-curve is flat
~64 on both shapes. (Pre-existing, unchanged: the 0.5 % idle-CPU gate line
trips on 32-vCPU shared-box arms on the BEFORE binary too — every matched
pair reads lower after.)

Gates (rebased on current main)

  • xerj-engine --lib: 731 passed / 0 failed
  • cargo fmt --check, clippy -D warnings (engine, all targets): clean
  • ES-YAML conformance, private port + throwaway dir: 1378 passed / 0 failed / 3 skipped

Not done here (tranche 2, issue stays open)

The remaining ~64 kB/idx: per-index DashMap shard-array laziness and a
quickwit-style idle-demotion loop (quickwit-ingest/src/ingest_v2/idle.rs:25-67,
models.rs:182-197) driving the existing #463 release path.

Refs #1024.

The #1024 idle-RSS audit reconstructed the live budget of one open index
at ~205 kB resident (4-vCPU idle-budget arm, flat in N; 183-194 kB/idx on
the 32-vCPU arm at N>=300). The single largest line item is the
request-cache seen-set: RequestCacheSeen::with_capacity(65_536) ran at
BOTH Index construction sites (the create and the open struct literals),
pre-writing a 131,072-bucket HashSet table plus a 512 KiB VecDeque ring -
measured 135 kB + ~1 kB resident per set under the server's jemalloc
config (background_thread:true, dirty_decay_ms:1000), i.e. ~136 of
~205 kB/idx, paid by every index at boot or create before it has ever
served a request_cache=true search. The fixture node's 14 system indices
pay it too.

Change (lazy allocation, no restructuring): RequestCacheSeen::new(cap)
starts both collections empty (HashSet::default() / VecDeque::default() -
zero bytes, zero written pages); record() grows them exactly as an
unsized HashSet/VecDeque grows, and the FIFO bound (evict oldest at cap)
is enforced exactly as before. First sight = miss, repeat = hit,
unchanged; the ES wire protocol is untouched. A busy index converges to
the same bounded footprint as before (65,536 entries, ~1.5 MiB logical);
an idle or never-request_cache-queried index holds nothing.

Design adapted (not copied) from quickwit's pay-per-use reader stack:
quickwit-search/src/leaf.rs:237-283 open_index_with_caches (Apache-2.0)
opens the whole per-split Index inside the leaf-search future and drops
it with the query, keeping idle memory at zero per split - the same
zero-until-used shape, applied to per-index bookkeeping.

Fail-before tests (crates/xerj-engine/src/index.rs,
request_cache_seen_idle_tests; ci-test profile, this tree):
- created_index_allocates_no_request_cache_seen_set FAILED before the
  change (HashSet capacity 131,072 at create), ok after
- reopened_index_allocates_no_request_cache_seen_set FAILED before (same
  on the boot/open path; also asserts the reopened index still serves its
  document, so laziness provably loses nothing), ok after
- track_request_cache_dedupes_and_evicts_when_lazily_grown pins the
  hit/miss contract and FIFO eviction at the cap through the public path
  (65,537 distinct hashes, oldest evicted and re-missed), ok before and
  after
Gates: xerj-engine --lib 730 passed / 0 failed; xerj-api suite all green;
ES-YAML conformance 1376 passed / 0 failed / 3 skipped (private port,
throwaway dir); rustfmt clean; clippy -D warnings clean.

Measured (idle-budget fixture, same host, back-to-back A/B on the
before-binary preserved from ed510a6, allocator pins
thp:never,dirty_decay_ms:0,muzzy_decay_ms:0):
- 4-vCPU taskset 0-3, N=450: 203.2 -> 68.0 kB/idx (-66%; RssAnon
  108,008 -> 45,028 kB). Idle CPU identical (0.1583% both arms),
  wakeups 25.3 -> 25.1/s, boot-to-green 6 ms both, 0 WAL-replay lines
  both.
- 32-vCPU, N=300: 186.4 -> 51.7 kB/idx (-72%; RssAnon 103,488 ->
  60,736 kB), boot 8 -> 7 ms, 0 WAL-replay lines both.
- The 32-vCPU after-run tripped the idle-CPU gate line (0.766% > 0.5%)
  with loadavg 28.99 at window start - foreign tenants on the shared
  unpinned box (the before-run read 0.475% at loadavg 5.39). A second
  attempt also landed on a loaded window (0.675%, loadavg 21.68; RSS
  54.3 kB/idx, consistent). The controlled comparison is the taskset'd
  4-vCPU pair, which read byte-identical idle CPU (0.1583% both arms)
  and wakeups 25.3 -> 25.1/s: this commit adds no timers or polling and
  strictly removes boot-time ctrl-page writes, so the unpinned-arm flap
  is box noise, not the change. Wakeups on the 32-vCPU arm stayed
  25.6 -> 28.0/s. Every other gate line passed on every run.
Issue #874's product line (204.8 kB/idx) is now met on both host shapes
with 3.0-4.0x margin.

Follow-ups (next tranches; the issue stays open): the remaining ~50-70
kB/idx is dominated by ~24 per-index DashMap 16-shard arrays (~2.1 kB
each, eagerly written at construction), WAL/memtable shard structs, and
3-4 parsed JSON copies of settings/schema/mapping; the cold-path endgame
is quickwit-style idle demotion (quickwit-ingest/src/ingest_v2/
idle.rs:25-67 CloseIdleShardsTask + models.rs:182-197 is_idle,
Apache-2.0 - one
last_write_instant timestamp and a weak-state periodic closer) driving
the existing #463 release path by idle age for cleanly-flushed indexes.

Refs #1024.
Appends the after-numbers to benchmarks/idle-budget/README.md's RSS
calibration section (history preserved; nothing rewritten) and commits the
result files they cite, verifying commit 6526e8ea (lazy request-cache
seen-set) on this host with the idle-budget fixture, ci-test binaries and
allocator pins on.

4-vCPU (taskset -c 0-3), per-index idle RSS (VmRSS minus empty node):
  N=150   before 204.9 kB  -> after  63.2 kB   (-69%)
  N=450   before 203.7 kB  -> after  63.3 kB   (-69%)
  N-slope (450-arm minus 150-arm)/300, VmRSS: 206.5 -> 63.7 kB/idx (-69%)
  (RssAnon slope: 205.7 -> 68.8). The subtraction and the slope now agree:
  the curve is flat at a genuinely O(N)-constant ~64 kB/idx, and issue
  #874's 0.2 MB/index product line is met with >3x margin. Idle CPU at
  N=450 0.225% -> 0.192%, wakeups 25.2 -> 21.4/s, boot 6 -> 6 ms, WAL
  replay 0 both.

32-vCPU (unpinned), N=450 — before arm added for this verification
(result-p1024-before-32cpu-n450.json; every earlier 32-vCPU reading was
N=300, so the shape had no N=450 baseline):
  per-index 199.0 kB -> 66.3 / 62.6 kB (two after runs; -67%)
  idle CPU 0.767% (loadavg 19) -> 0.650 / 0.633% (loadavg 3-4)
  wakeups 22.6 -> 25.9 / 22.2 per s
The >0.5% idle-CPU gate line trips on the BEFORE binary too on this arm
(0.767%): pre-existing for the unpinned many-core shape at N=450 on this
shared box, not a regression from 6526e8ea - every matched pair reads
lower after. The controlled 4-vCPU pair is the comparison; wakeups (the
timer-churn proxy) are flat on every arm.

Gates re-run for this verification: cargo fmt --check clean; clippy
--profile ci-test -p xerj-engine --all-targets -- -D warnings clean;
cargo test --profile ci-test -p xerj-engine 1221 passed / 0 failed /
0 ignored; ES-YAML conformance 1376 passed / 0 failed / 3 skipped
(private port 9630, throwaway data dir).

Refs #1024.
@cla-bot cla-bot Bot added the cla-signed label Sep 22, 2026
@xerj-org
xerj-org merged commit eec3459 into main Sep 22, 2026
22 of 23 checks passed
xerj-org added a commit that referenced this pull request Sep 26, 2026
Twelve PRs have merged since the rc.77 tag; the [Unreleased] section
carried only the #874 gauges entry. The release-notes gate
(.github/scripts/release-notes-gate.sh, issue #474) fails an rc.78
release PR whose notes do not cite every PR in the previous-tag..head
window, so each of these would have surfaced as a gate failure at cut
time instead of a review comment now.

Entries added (with PR links so the gate's coverage check resolves):
- Fixed: #1015 flush-drain freeze (PR #1018), #950 id-position maps +
  streamed reassembly (PR #1017), #1019 _delete_by_query paging (PR
  #1021), #1022 _update_by_query paging (PR #1023); the existing #874
  entry now cites its PR (#1020).
- Added: POST /{index}/_cache/clear (PR #1009), the systemone
  email-labelling benchmark answering discussion #1012 (PR #1026) with
  its measured numbers (1.000 templated / 0.625 at 0.902 confidence on
  the hard tier).
- Performance: request-cache seen-set lazy allocation, idle 206 -> 64
  kB/idx (PRs #1025, #1034).
- Documentation: README Jev section (PR #1010), the /_decide field
  report (PR #1027), llms.txt status catch-up (PR #1033, closing #1028).

Numbers are quoted only from the PR bodies' own verified runs.
Amakurai pushed a commit to Amakurai/xerj that referenced this pull request Sep 27, 2026
…, rc.78 queue

The 2026-09-21 review of this file was stale within hours: xerj-org#950 closed
19:33, xerj-org#941 19:38, xerj-org#1015 20:49 (all 2026-09-21), xerj-org#874 2026-09-22, and the
stage-1 gating trio xerj-org#937/xerj-org#938/xerj-org#939 closed the same evening — every one
recorded here as open or "under way". This pass re-checks every status
claim against live tracker state and rolls the file forward:

- Next release: the two "in flight" items (xerj-org#1015, xerj-org#950) are closed with
  fixes on main riding to rc.78 — the section now records what actually
  landed since the rc.77 tag (xerj-org#1009, xerj-org#1017, xerj-org#1018, xerj-org#1020, xerj-org#1021, xerj-org#1023,
  xerj-org#1025, xerj-org#1026, xerj-org#1033, xerj-org#1034) with PR links, including xerj-org#874's budget met
  at ~3x margin (206 -> 64 kB per idle index) and discussion xerj-org#1012's
  email-labelling measurement (1.000 templated tier / 0.625 at 0.902
  confidence on the hard tier — wrong-and-confident).
- Open defects: now xerj-org#1031 and xerj-org#1032 only; xerj-org#1015/xerj-org#950 removed (closed);
  trackers sentence corrected (xerj-org#941, xerj-org#874 closed — only xerj-org#298 remains by
  design).
- GA gate: "Close the CHANGELOG gap" marked closed 2026-09-26 (the
  rc.19-rc.70 backfill, PR xerj-org#1035); the record-complete note replaces the
  gap warning under "Shipping today".
- Zero-token: stage-1 gating trio recorded as fixed (xerj-org#991, xerj-org#995, xerj-org#979)
  with the old measured costs kept as the before-state; the xerj-org#940
  hash-seed spread annotated as a pre-fix measurement; stage-2 object
  storage marked done (xerj-org#965 wired in rc.77 — the old bullet contradicted
  this file's own "Shipping today" section); mail ingest "unreleased"
  corrected to "shipped in rc.75"; three stale "In flight:" labels on
  landed stage-1 items relabelled "History:".
- Mail-ingest memory line: xerj-org#948's runaway is fixed (xerj-org#1002) — the line
  now carries the fixed numbers and points at xerj-org#1032 for the residual.

Review line updated to 2026-09-26 (machine-checked format), scoped
honestly as a desk review. The milestones line is true again: rc.78
milestone created, all four open issues triaged (xerj-org#1030/xerj-org#1031/xerj-org#1032 ->
rc.78, xerj-org#298 -> GA).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant