Skip to content

fix(#1019): _delete_by_query pages the whole match set, ids-only - #1021

Merged
xerj-org merged 1 commit into
mainfrom
fix/issue-1019-delete-by-query-ids-only
Sep 22, 2026
Merged

xerj-org merged 1 commit into
mainfrom
fix/issue-1019-delete-by-query-ids-only

Conversation

@xerj-org

Copy link
Copy Markdown
Owner

What

Fixes #1019. Two defects, both in the by-query plumbing:

  1. Full _source materialization on a path that only needs ids — the delete search hydrated a full Hit { source: Value } per match (10,100 Value trees to read only hit.id); _source: false did not avoid it. This was the 6 GiB sort-stack residual the engine: RSS pinned above the memory watermark after bulk ingest into 1,526 indices — 14.8 GB anonymous memory for 1.2 GB on disk, not released at idle; node refuses every write until restart #950 fix recorded as follow-up headroom.
  2. size:10_000 single-shot cap — one call deleted at most 10,000 docs; total reported the truncated page length; batches was hardcoded 1. The autoindex client loops up to 1,000 passes as a workaround (each pass re-paying the full materializing search).

The fix, three layers: an internal SearchRequest::ids_only (serde(skip), hard-false in parse_request — cannot arrive from the wire) honored only when sort is _id/_doc-confined with no fields/highlight/aggs/etc.; keyset pagination (search_after on _id, default scroll_size 1000, ES-style max_docs); and a single-pass arm (Index::matching_ids_sorted) that reads only the #950 id-position maps + version map and never runs a search at all.

Root cause, two defects, both present in the HTTP runner
(es_compat.rs run_delete_by_query) and the engine loop
(Index::delete_by_query):

  1. ONE search, {"query":…,"size":10000,"from":0}, then delete whatever
    came back. Beyond 10 000 matching docs the call silently under-deleted
    while reporting success: total was the page length, batches was
    hardcoded 1. The autoindex client works around it by looping up to
    1 000 delete passes, each re-paying the search (esclient.rs:1658-1669).
  2. That search built a full Hit { source: Value } per match — 10 100
    hydrated Value trees per call to read nothing but hit.id.
    _source: false does not avoid this: the engine keeps the raw source
    for fields/highlight resolution, and the unsorted scan Value-parses
    every stored document anyway. This was the residual RSS transient
    engine: RSS pinned above the memory watermark after bulk ingest into 1,526 indices — 14.8 GB anonymous memory for 1.2 GB on disk, not released at idle; node refuses every write until restart #950 recorded as follow-up headroom.

Fail-before (both red on main, verbatim):

engine/crates/xerj-api/tests/delete_by_query_purges_entire_index.rs
delete_by_query_purges_past_ten_thousand_in_one_call
total must be the exact match count, not the page length:
{"total":10000,"deleted":10000,"batches":1,…}
left: Number(10000) right: Number(25000)
delete_by_query_honors_max_docs
with max_docs set, total is the processed count
left: Number(20) right: Number(7)
delete_by_query_reports_real_batch_count
10 docs at scroll_size 3 = 4 batches
left: Number(1) right: Number(4)
delete_by_query_paginates_selective_queries
20 matches at scroll_size 4 = 5 batches
left: Number(1) right: Number(5)
test result: FAILED. 0 passed; 4 failed

engine/crates/xerj-engine/tests/integration.rs
test_delete_by_query_purges_past_ten_thousand
every matching doc deleted, not just the first 10k page
left: 10000 right: 12000
test result: FAILED. 1 passed; 1 failed

Design.

  • Internal ids-only projection. SearchRequest::ids_only
    (xerj-query/src/ast.rs) is serde(skip) and hard-set false in
    parse_request, so it can never arrive from the wire (same contract as
    savings). search_inner honours it only when the request provably
    needs nothing beyond ids: sort non-empty and confined to _id/_doc,
    no fields/script_fields/highlight/explain/collapse/rescore/
    aggs/min_score. Every other shape runs exactly as before. At the
    three admission sites (stored-section scan + both hydration arms) the
    stored _id is extracted by a borrowed prefix scan
    (extract_stored_id_str; escape-bearing layouts fall back to the full
    parse, which still admits source-Null) and the hit carries
    source: Value::Null with the scorer skipped. The version-map liveness
    filter is kept verbatim. A precomputed cursor (id, ascending) declines
    docs at/below search_after with one memcmp BEFORE the liveness lookup
    and every allocation.

  • Keyset paged arm for arbitrary queries, modeled on the _reindex loop:
    flush the member first (paging and the id collection below are only
    correct over on-disk segments), then pull scroll_size pages sorted
    _id: asc with a search_after cursor. ES semantics: scroll_size
    default 1 000 clamped 1..=10_000, max_docs truncates processing and
    total then reports the processed count, batches is the real page
    count, total otherwise the exact page-1 match count,
    noops/version_conflicts 0, script resource failures fail-closed
    per batch before any id of that batch is deleted.

  • Single-pass arm for match_all/ids selectors. Measurement-driven:
    the paged arm alone re-scans every stored section once per page —
    O(N x pages). On the 252k-doc corpus the default-params purge measured
    43.5 s, 28.9 s after the cursor fast-decline; still quadratic, because
    every page must consider every doc. Index::matching_ids_sorted now
    collects the whole live match set in ONE pass over engine: RSS pinned above the memory watermark after bulk ingest into 1,526 indices — 14.8 GB anonymous memory for 1.2 GB on disk, not released at idle; node refuses every write until restart #950's cached
    id_pos_map_for id maps plus the version map (one String per live
    match, ~48 B), sorts it, and the runner deletes it in scroll_size
    batches. Observable semantics unchanged (batches = ceil(n/scroll_size));
    like ES's by-query, the match set is the start-of-run snapshot and a
    concurrently-vanished id is tolerated. Any other query shape (or a
    segment without a complete id index) takes the paged arm; None falls
    back, never errors.

Reference-code retrieval, honestly: xc.py was run this session and
FAILED for this task's domains — the quickwit and elasticsearch corpora
are not in the sandbox index, and building them was outside this task's
hard constraints (the shared corpus server/index must not be touched).
The design is XERJ's own. The one adapted approach — resolving document
ids from an id map instead of the document store — was recorded at
commit d1e8056 (#950): meilisearch external_documents_ids()
(crates/meilisearch/src/routes/indexes/documents.rs:2284, MIT, approach
only) and quickwit warm_up_terms (quickwit-search/src/leaf.rs:475,
Apache-2.0); those file:line references are from the prior session and
approximate. #1019 builds on that #950 machinery. No Elasticsearch (AGPL/
SSPL/Elastic) or sonic (GPL) code was read for or copied into this change;
ES was consulted only for wire semantics (scroll_size, max_docs,
total, batches).

Measurement — synthetic corpus, this session, private ports + throwaway
data dirs, wall via curl time_total, RSS via /proc//status sampled
at 250 ms: 252 000 docs / 74.5 MB NDJSON ingested into one index
(16 segments). before = main 62fa237 release binary, after = this
branch's release binary.

single _delete_by_query, default params (match_all):
before wall 0.54 s deleted 10 000 / 252 000 (defect 1)
peak RSS 1 234 216 kB VmHWM 1 256 268 kB
before, client workaround loop until _count==0
wall 3.78 s deleted 252 000
peak RSS 1 285 416 kB VmHWM 1 354 104 kB
after wall 1.10 s deleted 252 000 / 252 000
total 252 000 batches 252 failures []
peak RSS 466 100 kB VmHWM 471 764 kB

3.4x faster than the pre-fix workaround loop AND complete, with ~62%
lower peak RSS. Intermediates on the same corpus (all this session):
ids-only + keyset paging alone 43.5 s default / 9.1 s at scroll 10k;

Gates (this branch, final code): cargo fmt --check clean; clippy
-p xerj-engine -p xerj-api -p xerj-query --all-targets clean; full
xerj-engine suite green; full xerj-api suite green (68 test binaries);
ES-YAML conformance suite 1378 passed / 0 failed / 3 skipped against a
private port with a throwaway data dir (main: 1376/0/3; +2 new
delete_by_query cases). New coverage: 4 HTTP-level tests, 1 engine
integration test, 1 engine lib test pinning source: Null + gap-free
dupe-free keyset pages, 2 ES-YAML cases, all fail-before on main.

Fixes #1019.


Independent verification (second agent, nothing trusted from the implementer): fail-before reproduced by reverting the four fix files to main (API 0/4 passed — total 10000 vs 25000; engine 10000 vs 12000); all new tests green on the branch; ES-YAML 1378 passed / 0 failed / 3 skipped (main is 1376/0/3 — the +2 are new delete_by_query pinning cases), run against a release binary rebuilt from the committed tree and md5-verified identical to the measured one; fmt + clippy clean, zero warnings; wire-level check: 15,000-doc _bulk then one default _delete_by_query purged all 15,000 in 118 ms.

Adversarial diff review confirmed: no _source on the delete path (pinned by test: source: Value::Null + gap-free/dupe-free keyset pages), pagination runs to exhaustion with a 10M backstop, error paths fail closed (malformed query 400s before any deletion; failed flush aborts; script-limit refusal checked per batch before that batch's deletes), and the one deliberate O(N) point (match set as Vec<String>, ~12 MB at 252k) is documented and bounded.

Environmental note: re-running the ES-YAML gate on this sandbox needed flood_stage=99% — disk sits at 95%, tripping the flood-stage write block by design. Not a branch defect.

Root cause, two defects, both present in the HTTP runner
(es_compat.rs run_delete_by_query) and the engine loop
(Index::delete_by_query):

1. ONE search, `{"query":…,"size":10000,"from":0}`, then delete whatever
   came back. Beyond 10 000 matching docs the call silently under-deleted
   while reporting success: `total` was the page length, `batches` was
   hardcoded 1. The autoindex client works around it by looping up to
   1 000 delete passes, each re-paying the search (esclient.rs:1658-1669).
2. That search built a full `Hit { source: Value }` per match — 10 100
   hydrated `Value` trees per call to read nothing but `hit.id`.
   `_source: false` does not avoid this: the engine keeps the raw source
   for `fields`/highlight resolution, and the unsorted scan Value-parses
   every stored document anyway. This was the residual RSS transient
   #950 recorded as follow-up headroom.

Fail-before (both red on main, verbatim):

  engine/crates/xerj-api/tests/delete_by_query_purges_entire_index.rs
    delete_by_query_purges_past_ten_thousand_in_one_call
      total must be the exact match count, not the page length:
      {"total":10000,"deleted":10000,"batches":1,…}
        left: Number(10000)  right: Number(25000)
    delete_by_query_honors_max_docs
      with max_docs set, total is the processed count
        left: Number(20)  right: Number(7)
    delete_by_query_reports_real_batch_count
      10 docs at scroll_size 3 = 4 batches
        left: Number(1)  right: Number(4)
    delete_by_query_paginates_selective_queries
      20 matches at scroll_size 4 = 5 batches
        left: Number(1)  right: Number(5)
    test result: FAILED. 0 passed; 4 failed

  engine/crates/xerj-engine/tests/integration.rs
    test_delete_by_query_purges_past_ten_thousand
      every matching doc deleted, not just the first 10k page
        left: 10000  right: 12000
    test result: FAILED. 1 passed; 1 failed

Design.

* Internal ids-only projection. `SearchRequest::ids_only`
  (xerj-query/src/ast.rs) is `serde(skip)` and hard-set `false` in
  `parse_request`, so it can never arrive from the wire (same contract as
  `savings`). `search_inner` honours it only when the request provably
  needs nothing beyond ids: sort non-empty and confined to `_id`/`_doc`,
  no `fields`/`script_fields`/`highlight`/`explain`/`collapse`/`rescore`/
  `aggs`/`min_score`. Every other shape runs exactly as before. At the
  three admission sites (stored-section scan + both hydration arms) the
  stored `_id` is extracted by a borrowed prefix scan
  (`extract_stored_id_str`; escape-bearing layouts fall back to the full
  parse, which still admits source-Null) and the hit carries
  `source: Value::Null` with the scorer skipped. The version-map liveness
  filter is kept verbatim. A precomputed cursor `(id, ascending)` declines
  docs at/below `search_after` with one memcmp BEFORE the liveness lookup
  and every allocation.

* Keyset paged arm for arbitrary queries, modeled on the `_reindex` loop:
  flush the member first (paging and the id collection below are only
  correct over on-disk segments), then pull `scroll_size` pages sorted
  `_id: asc` with a `search_after` cursor. ES semantics: `scroll_size`
  default 1 000 clamped 1..=10_000, `max_docs` truncates processing and
  `total` then reports the processed count, `batches` is the real page
  count, `total` otherwise the exact page-1 match count,
  `noops`/`version_conflicts` 0, script resource failures fail-closed
  per batch before any id of that batch is deleted.

* Single-pass arm for `match_all`/`ids` selectors. Measurement-driven:
  the paged arm alone re-scans every stored section once per page —
  O(N x pages). On the 252k-doc corpus the default-params purge measured
  43.5 s, 28.9 s after the cursor fast-decline; still quadratic, because
  every page must consider every doc. `Index::matching_ids_sorted` now
  collects the whole live match set in ONE pass over #950's cached
  `id_pos_map_for` id maps plus the version map (one `String` per live
  match, ~48 B), sorts it, and the runner deletes it in `scroll_size`
  batches. Observable semantics unchanged (`batches` = ceil(n/scroll_size));
  like ES's by-query, the match set is the start-of-run snapshot and a
  concurrently-vanished id is tolerated. Any other query shape (or a
  segment without a complete id index) takes the paged arm; `None` falls
  back, never errors.

Reference-code retrieval, honestly: `xc.py` was run this session and
FAILED for this task's domains — the quickwit and elasticsearch corpora
are not in the sandbox index, and building them was outside this task's
hard constraints (the shared corpus server/index must not be touched).
The design is XERJ's own. The one adapted approach — resolving document
ids from an id map instead of the document store — was recorded at
commit d1e8056 (#950): meilisearch `external_documents_ids()`
(crates/meilisearch/src/routes/indexes/documents.rs:2284, MIT, approach
only) and quickwit `warm_up_terms` (quickwit-search/src/leaf.rs:475,
Apache-2.0); those file:line references are from the prior session and
approximate. #1019 builds on that #950 machinery. No Elasticsearch (AGPL/
SSPL/Elastic) or sonic (GPL) code was read for or copied into this change;
ES was consulted only for wire semantics (`scroll_size`, `max_docs`,
`total`, `batches`).

Measurement — synthetic corpus, this session, private ports + throwaway
data dirs, wall via curl time_total, RSS via /proc/<pid>/status sampled
at 250 ms: 252 000 docs / 74.5 MB NDJSON ingested into one index
(16 segments). before = main 62fa237 release binary, after = this
branch's release binary.

  single _delete_by_query, default params (match_all):
    before        wall 0.54 s   deleted 10 000 / 252 000  (defect 1)
                 peak RSS 1 234 216 kB   VmHWM 1 256 268 kB
    before, client workaround loop until _count==0
                 wall 3.78 s   deleted 252 000
                 peak RSS 1 285 416 kB   VmHWM 1 354 104 kB
    after         wall 1.10 s   deleted 252 000 / 252 000
                 total 252 000   batches 252   failures []
                 peak RSS   466 100 kB   VmHWM   471 764 kB

  3.4x faster than the pre-fix workaround loop AND complete, with ~62%
  lower peak RSS. Intermediates on the same corpus (all this session):
  ids-only + keyset paging alone 43.5 s default / 9.1 s at scroll 10k;
  + cursor fast-decline 28.9 s / 7.4 s; + single-pass arm 1.10 s. The
  ~900 MB RSS transient common to both before runs is NOT the delete
  search itself: a single plain wire-level sorted search returned having
  added ~220 MB and grew to HWM 1.21 GB three seconds AFTER the response
  — background post-search merge/cache hydration over the 16-segment
  corpus (the #950 follow-up area, out of #1019's scope). The single-pass
  arm never runs that search, so its peak is the delete path's own.

Gates (this branch, final code): cargo fmt --check clean; clippy
-p xerj-engine -p xerj-api -p xerj-query --all-targets clean; full
xerj-engine suite green; full xerj-api suite green (68 test binaries);
ES-YAML conformance suite 1378 passed / 0 failed / 3 skipped against a
private port with a throwaway data dir (main: 1376/0/3; +2 new
delete_by_query cases). New coverage: 4 HTTP-level tests, 1 engine
integration test, 1 engine lib test pinning `source: Null` + gap-free
dupe-free keyset pages, 2 ES-YAML cases, all fail-before on main.

Fixes #1019.
@cla-bot cla-bot Bot added the cla-signed label Sep 21, 2026
@xerj-org
xerj-org merged commit 7f0a085 into main Sep 22, 2026
22 checks passed
@xerj-org
xerj-org deleted the fix/issue-1019-delete-by-query-ids-only branch September 22, 2026 02:51
xerj-org added a commit that referenced this pull request Sep 22, 2026
…, WILL BE REVERTED

PR #1023 went merge-dirty against main (es_compat.rs changed on both sides
when #1009/#1021 landed), so pushes to the branch stopped creating
pull_request CI runs entirely - GitHub cannot build refs/pull/1023/merge and
silently creates no run (observed: two pushes, only the CLA and code-scanning
workflows fired). The diagnostic run therefore cannot reach Build + Test the
normal way.

Two temporary triggers, both on the branch ref (our exact tree, no merge):

- diag1023.yml: minimal capture job replicating Build + Test's runner,
  toolchain pin, and rust-cache key - runs the xerj-api lib suite at
  --test-threads=2, then 40 isolated --nocapture iterations of the target
  test. [DIAG2] eprintln fires from tokio worker threads, so it reaches the
  job log whether the test passes or fails.
- ci.yml on:workflow_dispatch: so the full authoritative Build + Test shape
  can also be dispatched on the branch. PR-context steps are if:-gated on
  pull_request and simply skip.

Both are reverted with the instrumentation.
xerj-org added a commit that referenced this pull request Sep 26, 2026
Twelve PRs have merged since the rc.77 tag; the [Unreleased] section
carried only the #874 gauges entry. The release-notes gate
(.github/scripts/release-notes-gate.sh, issue #474) fails an rc.78
release PR whose notes do not cite every PR in the previous-tag..head
window, so each of these would have surfaced as a gate failure at cut
time instead of a review comment now.

Entries added (with PR links so the gate's coverage check resolves):
- Fixed: #1015 flush-drain freeze (PR #1018), #950 id-position maps +
  streamed reassembly (PR #1017), #1019 _delete_by_query paging (PR
  #1021), #1022 _update_by_query paging (PR #1023); the existing #874
  entry now cites its PR (#1020).
- Added: POST /{index}/_cache/clear (PR #1009), the systemone
  email-labelling benchmark answering discussion #1012 (PR #1026) with
  its measured numbers (1.000 templated / 0.625 at 0.902 confidence on
  the hard tier).
- Performance: request-cache seen-set lazy allocation, idle 206 -> 64
  kB/idx (PRs #1025, #1034).
- Documentation: README Jev section (PR #1010), the /_decide field
  report (PR #1027), llms.txt status catch-up (PR #1033, closing #1028).

Numbers are quoted only from the PR bodies' own verified runs.
Amakurai pushed a commit to Amakurai/xerj that referenced this pull request Sep 27, 2026
…ext/enrich counts

xerj-org#1019's verify pass found three more one-shot `size:10000` truncations.
All three silently did a third (or less) of the requested work while
reporting success; this commit pages or folds all three to completion.

Root cause, three sites, same shape:

1. `_update_by_query` (es_compat.rs run_update_by_query): ONE
   `{"query":…,"size":10000,"from":0}` search, update whatever came
   back, `total` was the page length, `batches` hardcoded 1. Past
   10 000 matches the call silently under-updated. `max_docs` and
   `scroll_size` were not implemented at all.

2. significant_text profiler fast path (compute_sig_text_debug): the
   profile `debug` counters were computed from the first 10 000
   hydrated hits, so `extract_count`/`values_fetched`/`chars_fetched`/
   `collect_analyzed_count` froze at 10 000 on any larger corpus.

3. enrich policy `_execute`: same one-shot copy — the `.enrich-*`
   index was materialised from only the first 10 000 source docs and
   `records` reported the page length.

Fail-before (red on the starting tree 64c698e, verbatim):

  xerj-api/tests/update_by_query_updates_entire_index.rs  0 passed; 6 failed
    update_by_query_scripts_past_ten_thousand_in_one_call
      total must be the exact match count, not the page length:
      {"total":10000,"updated":10000,"batches":1,…}
        left: Number(10000)  right: Number(25000)
    update_by_query_reindexes_past_ten_thousand_in_one_call
        left: Number(10000)  right: Number(25000)
    update_by_query_honors_max_docs
      with max_docs set, total is the processed count
        left: Number(20)  right: Number(7)
    update_by_query_reports_real_batch_count
      10 docs at scroll_size 3 = 4 batches
        left: Number(1)  right: Number(4)
    update_by_query_paginates_selective_queries
      20 matches at scroll_size 4 = 5 batches
        left: Number(1)  right: Number(5)
    update_by_query_applies_script_exactly_once_per_doc_when_paged
      2500 docs at scroll_size 100 = 25 batches
        left: Number(1)  right: Number(25)

  xerj-api/tests/significant_text_profile_counts_every_doc.rs  2 passed; 1 failed
    significant_text_profile_counts_past_ten_thousand
      extract_count must count every matched doc:
      {"values_fetched":10000,"chars_fetched":110000,
       "extract_count":10000,"collect_analyzed_count":20000}
        left: Number(10000)  right: Number(10500)

  xerj-api/tests/enrich_execute_copies_entire_index.rs  1 passed; 1 failed
    enrich_execute_copies_past_ten_thousand
      records must be the exact source count, not the page length:
      {"enrich_index":".enrich-pol-big","records":10000}
        left: Number(10000)  right: Number(10500)

Design — all three reuse xerj-org#1019's machinery; nothing external.

* SITE 1, run_update_by_query rewritten on the delete runner's shape
  (es_compat.rs:26792): flush precondition; the query is validated once
  through the same `{"query":…,"size":scroll_size,"sort":[{"_id":"asc"}]}`
  body; `scroll_size` defaults 1000 clamped 1..=10_000, `max_docs`
  truncates processing and `total` then reports the processed count,
  `batches` is the real page count. The 30 s SCRIPTED_UPDATE_BUDGET
  deadline and fault capture now span the whole paged loop (the runner
  is Box::pin'd — the inline state machine overflowed the 2 MiB
  debug-profile test stack). Script mode + match_all/ids selectors take
  xerj-org#1019's single-pass `matching_ids_sorted` arm (one O(N) id
  materialisation, transformed per id in scroll_size chunks);
  everything else takes `_id`-keyset `search_after` pages with the
  ids_only projection (xerj-org#1019). Per-id work goes through
  `transform_document_serialized` inside `transform_one`
  (es_compat.rs:27132): a doc that vanished mid-run is skipped
  silently, an update error lands in `failures` per ES semantics.
  The no-script arm (reindex-in-place, pipelines) is preserved verbatim
  but paged with `_source:true` pages.

  Mid-run rewrite hazard, unique to update: a batch that mutates docs
  repopulates the memtable, and `_id`-keyset paging is only correct
  over flushed segments (the invariant documented at the reindex loop
  and run_delete_by_query) — an unflushed next page could resurface a
  doc at/below the cursor and apply the script TWICE. So every batch
  that mutated >= 1 doc flushes the index before the next page is
  searched; a failed flush aborts the run rather than risk it. The
  2500-doc scroll_size-100 regression test pins exactly-once semantics.

* SITE 2, exact counting without materialisation: the significant_text
  debug counters are now an incremental accumulator
  (SigTextAccumulator, es_compat.rs:24897 — sums and per-group token
  sets, i.e. order-independent) fed by a new single-pass engine walker
  `Index::fold_match_sources` (index.rs:17891): memtable walk plus
  per-segment stored-section decode under the existing
  decoded_stored_cache/stored_slices_cache single-flight locks, with
  the version-map liveness filter, tombstone/ghost-window handling and
  seen-set dedup, matching every document with the same resolution
  pipeline and `doc_matches_query_typed` matcher the memtable scan arm
  of search uses (xerj-org#396: buffered answers must not change at flush).
  A fold, not keyset paging: a read endpoint cannot flush (that would
  be a write side effect from _search), and without a flush the
  keyset-page invariant is unavailable — order-independent counters
  need no ordering at all, so one pass over a consistent segment
  snapshot is exact. The debug json shape is unchanged (pinned by the
  small-corpus test and the existing significant_text.yml pins).

* SITE 3, enrich `_execute` adopts the `_reindex` keyset loop
  (es_compat.rs:32344): flush each source, then 1 000-doc pages sorted
  `_id: asc` with a `search_after` cursor and `_source:true`; first
  write error fails the request; `records` is the exact materialised
  count; response shape unchanged. No new wire params.

Measurement — synthetic corpus (30 000 generated docs, ~250 B each,
stated as synthetic), this session, private ports 9622/9623 +
throwaway data dirs, wall via curl time_total, RSS via
/proc/<pid>/status sampled at 250 ms. before = base 64c698e release
binary, after = this branch's release binary:

  _update_by_query, match_all + `ctx._source.touched = 1`:
                          before            after
    total / updated       10 000 / 30 000   30 000 / 30 000
    batches               1 (hardcoded)     30 (real, scroll 1000)
    wall                  0.24 s            1.00 s
    touched docs (_count) 10 000            30 000
    last doc in _id order no touched field   touched = 1
    peak RSS (window)     384 MB            554 MB
    timed_out             false             false

  enrich _execute (one source index):
    records / .enrich-*    10 000 / 30 000   30 000 / 30 000
    wall                   0.24 s            2.89 s

  significant_text profile counters (match_all over the body field):
    extract_count          10 000            30 000
    chars_fetched          1 800 000         5 400 000
    collect_analyzed_count 180 000           540 000
    wall                   0.34 s            0.33 s

  before is "faster" because it does a third of the work and reports
  the page length as the total. Per-doc amortised update cost moved
  24 -> 33 us/doc for completeness' extra writes + 30 flushes. The
  sig-text fold counts 3x the documents in the same wall time because
  it replaces the search+hydrate path entirely.

Gates (final tree): cargo fmt --check clean; clippy -p xerj-api
-p xerj-engine --all-targets 0 warnings; cargo test --profile ci-test
-p xerj-api -p xerj-engine 146 test binaries ok, 1898 passed, 0
failed (ci-test is CI's profile; the dev-profile run additionally
passes all xerj-api binaries but aborts in the pre-existing
hybrid_fused_order_is_stable debug stack overflow, verified present
on the starting tree 64c698e before any of this branch's changes);
ES-YAML conformance 1380 passed / 0 failed / 3 skipped against a
private port with a throwaway data dir (starting tree: 1378/0/3;
+2 new update_by_query cases). New coverage: 6 update_by_query HTTP
tests, 2 enrich HTTP tests, 3 sig-text HTTP tests, 2 ES-YAML cases —
all fail-before on the starting tree.

Reference-code: none retrieved — the adapted pattern is XERJ's own,
established in-tree by xerj-org#1019 (d1e8056 lineage); CLAUDE.md's
retrieval mandate does not bind on changes confined to this
repository's own code. No Elasticsearch (AGPL/SSPL/Elastic) or sonic
(GPL) code was read or copied; ES was consulted only for wire
semantics (max_docs, scroll_size, total, batches).

stacked on fix/issue-1019-delete-by-query-ids-only (xerj-org#1021); merge that first

Fixes xerj-org#1022.
Amakurai pushed a commit to Amakurai/xerj that referenced this pull request Sep 27, 2026
…, rc.78 queue

The 2026-09-21 review of this file was stale within hours: xerj-org#950 closed
19:33, xerj-org#941 19:38, xerj-org#1015 20:49 (all 2026-09-21), xerj-org#874 2026-09-22, and the
stage-1 gating trio xerj-org#937/xerj-org#938/xerj-org#939 closed the same evening — every one
recorded here as open or "under way". This pass re-checks every status
claim against live tracker state and rolls the file forward:

- Next release: the two "in flight" items (xerj-org#1015, xerj-org#950) are closed with
  fixes on main riding to rc.78 — the section now records what actually
  landed since the rc.77 tag (xerj-org#1009, xerj-org#1017, xerj-org#1018, xerj-org#1020, xerj-org#1021, xerj-org#1023,
  xerj-org#1025, xerj-org#1026, xerj-org#1033, xerj-org#1034) with PR links, including xerj-org#874's budget met
  at ~3x margin (206 -> 64 kB per idle index) and discussion xerj-org#1012's
  email-labelling measurement (1.000 templated tier / 0.625 at 0.902
  confidence on the hard tier — wrong-and-confident).
- Open defects: now xerj-org#1031 and xerj-org#1032 only; xerj-org#1015/xerj-org#950 removed (closed);
  trackers sentence corrected (xerj-org#941, xerj-org#874 closed — only xerj-org#298 remains by
  design).
- GA gate: "Close the CHANGELOG gap" marked closed 2026-09-26 (the
  rc.19-rc.70 backfill, PR xerj-org#1035); the record-complete note replaces the
  gap warning under "Shipping today".
- Zero-token: stage-1 gating trio recorded as fixed (xerj-org#991, xerj-org#995, xerj-org#979)
  with the old measured costs kept as the before-state; the xerj-org#940
  hash-seed spread annotated as a pre-fix measurement; stage-2 object
  storage marked done (xerj-org#965 wired in rc.77 — the old bullet contradicted
  this file's own "Shipping today" section); mail ingest "unreleased"
  corrected to "shipped in rc.75"; three stale "In flight:" labels on
  landed stage-1 items relabelled "History:".
- Mail-ingest memory line: xerj-org#948's runaway is fixed (xerj-org#1002) — the line
  now carries the fixed numbers and points at xerj-org#1032 for the residual.

Review line updated to 2026-09-26 (machine-checked format), scoped
honestly as a desk review. The milestones line is true again: rc.78
milestone created, all four open issues triaged (xerj-org#1030/xerj-org#1031/xerj-org#1032 ->
rc.78, xerj-org#298 -> GA).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

delete_by_query materializes every hit _source and caps at size:10_000 — 6 GiB sort stacks, ~30-min search phase on 252k docs

1 participant