Skip to content

perf(extract, features): use every thread when the batch is small - #106

Merged
RobbinBouwmeester merged 6 commits into
mainfrom
perf/extract-window-parallel
Sep 21, 2026
Merged

RobbinBouwmeester merged 6 commits into
mainfrom
perf/extract-window-parallel

Conversation

@RobbinBouwmeester

Copy link
Copy Markdown
Member

Summary

Two stages left most of their threads idle, and the isolation-window-group search (#105) makes that worse rather than better: a group has few windows, and the window was the unit of parallelism. Measured on one band of the 8-12-mer immunopeptidomics run with --threads 24: extract 3.2-3.7 cores busy for the whole stage, features 1.6-2.8.

  • extract parallelised across the windows in flight, so windows_in_flight: 4 (set low because the hit accumulator holds those windows, and that is the stage's memory) capped it at four working threads. Each window is now also split across its candidate range. A candidate belongs to exactly one sub-range, so the sub-results are disjoint: they hold the window's hits once between them (no extra memory, which a scan-axis split would have cost), they concatenate instead of merging per candidate, and each candidate's hits still arrive in ascending scan order, so every float reduction downstream is unchanged. The candidate axis is also where the work is, since a peak's cost is dominated by walking its posting list. windows_in_flight becomes a pure memory knob.
  • features decoded each chromatogram chunk single-threaded and then computed on it, in series; the decode dominated (3.7 of a 4-minute stage on one band). The loader now runs on its own thread one chunk ahead, so decode and compute overlap, with at most two chunks resident.

Output is unchanged

  • CI smoke peptides.tsv / proteins.tsv hashes unchanged; all five stage artifacts byte-identical to the previous binary's on the fixture.
  • Real scale, one 16.5M-precursor band of the immunopeptidomics run: psms_extracted, chromatograms and features byte-identical to the old binary's, at 12.8 min against 39.6 min. The two runs had different machine load (the old one shared the host with eleven other extracts), so treat the ratio as indicative, not as a measured speed-up.
  • New test candidate_range_split_reproduces_the_unsplit_accumulation: same candidates, same hits, same per-candidate hit order whether a window is probed whole or in three sub-ranges.

Notes

  • accumulate_groups spawns its producer into the rayon pool and consumes on the calling thread. The engine calls it from the main thread, which is never a pool worker (build_global), so a one-thread engine is fine; a test that calls it under pool.install with one thread would deadlock, and the test comment says so.
  • The two-pass co-elution path (extract_twopass_windows, used by the peak_claim co-elution modes) still parallelises across windows only. Same treatment is possible there and is not in this PR.
  • Stacks cleanly with feat(run): search a run one isolation-window group at a time (groups.window_groups) #105: independent files apart from config.rs documentation.

Validation

cargo fmt --check, cargo clippy --workspace --all-targets -- -D warnings, cargo test --workspace (291 tests), ci/smoke.sh with unchanged hashes.

🤖 Generated with Claude Code

RobbinBouwmeester and others added 6 commits September 21, 2026 15:47
Two stages left most of their threads idle, which grouping made worse rather than
better: a window group has few isolation windows, and that was the unit of parallelism.
Measured on one band of the 8-12-mer immunopeptidomics run with `--threads 24`: extract
3.2-3.7 cores busy, features 1.6-2.8.

extract parallelised across the windows in flight, so `windows_in_flight: 4` (set low
because the hit accumulator holds those windows, and that is the stage's memory) capped it
at four threads. Each window is now also split across its CANDIDATE range: a candidate
belongs to exactly one sub-range, so the sub-results are disjoint, they hold the window's
hits once between them (no extra memory, which a scan-axis split would have cost), they
concatenate instead of merging per candidate, and each candidate's hits still arrive in
ascending scan order, so every float reduction downstream is unchanged. Splitting the
candidate axis is also where the work is: a peak's cost is dominated by walking its posting
list, which the narrowed sub-range divides. `windows_in_flight` becomes a memory knob only.

features decoded each chromatogram chunk single-threaded and then computed on it, in
series; the decode dominated (3.7 of a 4-minute stage on one band). The loader now runs on
its own thread one chunk ahead, so decode and compute overlap, with at most two chunks
resident. Each chunk travels with the fragment-name table as it stood when the chunk was
read, which is what the serial code saw.

Byte-identical output: the CI smoke's `peptides.tsv`/`proteins.tsv` hashes are unchanged
and all five stage artifacts match the previous binary's on the fixture; at real scale, one
16.5M-precursor band reproduced the old binary's `psms_extracted`, `chromatograms` and
`features` exactly, in 12.8 min against 39.6 (the two runs had different machine load, so
treat the ratio as indicative). A new test asserts the split is a partition: the
accumulator, including each candidate's hit order, is the same whether a window is probed
whole or in three sub-ranges.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The parallelism change reworded that field's documentation; the generated reference and
schema are derived from it and CI checks they are current.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Two extracts of one band at 1 and at 4 windows in flight hold the same 29,028,466
chromatogram rows in the same order, with every per-column sum equal over 10.4 billion
trace elements, and still differ byte for byte: the parquet row group boundaries follow
the flush batches, which follow the batch size. Measured while checking the parallelism
change, where a byte comparison across settings looked like a regression and was not one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…arallel

# Conflicts:
#	configs/config-schema.json
#	docs/24_config_reference.md
…pOmics/MuMDIA into perf/extract-window-parallel

# Conflicts:
#	configs/config-schema.json
#	docs/24_config_reference.md
@RobbinBouwmeester
RobbinBouwmeester merged commit 0feed7d into main Sep 21, 2026
12 checks passed
@RobbinBouwmeester RobbinBouwmeester mentioned this pull request Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant