perf(extract, features): use every thread when the batch is small - #106
Merged
Merged
Conversation
Two stages left most of their threads idle, which grouping made worse rather than better: a window group has few isolation windows, and that was the unit of parallelism. Measured on one band of the 8-12-mer immunopeptidomics run with `--threads 24`: extract 3.2-3.7 cores busy, features 1.6-2.8. extract parallelised across the windows in flight, so `windows_in_flight: 4` (set low because the hit accumulator holds those windows, and that is the stage's memory) capped it at four threads. Each window is now also split across its CANDIDATE range: a candidate belongs to exactly one sub-range, so the sub-results are disjoint, they hold the window's hits once between them (no extra memory, which a scan-axis split would have cost), they concatenate instead of merging per candidate, and each candidate's hits still arrive in ascending scan order, so every float reduction downstream is unchanged. Splitting the candidate axis is also where the work is: a peak's cost is dominated by walking its posting list, which the narrowed sub-range divides. `windows_in_flight` becomes a memory knob only. features decoded each chromatogram chunk single-threaded and then computed on it, in series; the decode dominated (3.7 of a 4-minute stage on one band). The loader now runs on its own thread one chunk ahead, so decode and compute overlap, with at most two chunks resident. Each chunk travels with the fragment-name table as it stood when the chunk was read, which is what the serial code saw. Byte-identical output: the CI smoke's `peptides.tsv`/`proteins.tsv` hashes are unchanged and all five stage artifacts match the previous binary's on the fixture; at real scale, one 16.5M-precursor band reproduced the old binary's `psms_extracted`, `chromatograms` and `features` exactly, in 12.8 min against 39.6 (the two runs had different machine load, so treat the ratio as indicative). A new test asserts the split is a partition: the accumulator, including each candidate's hit order, is the same whether a window is probed whole or in three sub-ranges. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The parallelism change reworded that field's documentation; the generated reference and schema are derived from it and CI checks they are current. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Two extracts of one band at 1 and at 4 windows in flight hold the same 29,028,466 chromatogram rows in the same order, with every per-column sum equal over 10.4 billion trace elements, and still differ byte for byte: the parquet row group boundaries follow the flush batches, which follow the batch size. Measured while checking the parallelism change, where a byte comparison across settings looked like a regression and was not one. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…arallel # Conflicts: # configs/config-schema.json # docs/24_config_reference.md
…pOmics/MuMDIA into perf/extract-window-parallel # Conflicts: # configs/config-schema.json # docs/24_config_reference.md
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two stages left most of their threads idle, and the isolation-window-group search (#105) makes that worse rather than better: a group has few windows, and the window was the unit of parallelism. Measured on one band of the 8-12-mer immunopeptidomics run with
--threads 24: extract 3.2-3.7 cores busy for the whole stage, features 1.6-2.8.windows_in_flight: 4(set low because the hit accumulator holds those windows, and that is the stage's memory) capped it at four working threads. Each window is now also split across its candidate range. A candidate belongs to exactly one sub-range, so the sub-results are disjoint: they hold the window's hits once between them (no extra memory, which a scan-axis split would have cost), they concatenate instead of merging per candidate, and each candidate's hits still arrive in ascending scan order, so every float reduction downstream is unchanged. The candidate axis is also where the work is, since a peak's cost is dominated by walking its posting list.windows_in_flightbecomes a pure memory knob.Output is unchanged
peptides.tsv/proteins.tsvhashes unchanged; all five stage artifacts byte-identical to the previous binary's on the fixture.psms_extracted,chromatogramsandfeaturesbyte-identical to the old binary's, at 12.8 min against 39.6 min. The two runs had different machine load (the old one shared the host with eleven other extracts), so treat the ratio as indicative, not as a measured speed-up.candidate_range_split_reproduces_the_unsplit_accumulation: same candidates, same hits, same per-candidate hit order whether a window is probed whole or in three sub-ranges.Notes
accumulate_groupsspawns its producer into the rayon pool and consumes on the calling thread. The engine calls it from the main thread, which is never a pool worker (build_global), so a one-thread engine is fine; a test that calls it underpool.installwith one thread would deadlock, and the test comment says so.extract_twopass_windows, used by thepeak_claimco-elution modes) still parallelises across windows only. Same treatment is possible there and is not in this PR.config.rsdocumentation.Validation
cargo fmt --check,cargo clippy --workspace --all-targets -- -D warnings,cargo test --workspace(291 tests),ci/smoke.shwith unchanged hashes.🤖 Generated with Claude Code