Repository navigation
perf(tableau): optimize measurement for generalized tableau - #56
Conversation
|
There was a problem hiding this comment.
Pull request overview
This PR optimizes GeneralizedTableau::measure() (the dominant runtime hotspot for generalized tableau simulations with multiple T gates) by reducing allocations and HashMap work, and by simplifying overlap computation to accumulate only the real component.
Changes:
- Reworks
GeneralizedTableau::measure()to avoid cloning coefficients, add a case-b fast path, and optimize case-a partitioning/drain behavior. - Removes
trim_coefficients_for_measurement()(now inlined/handled inmeasure()’s case-b path). - Adds a new Criterion benchmark (
measure-scaling) and an ad-hoc profiling example, plus autotune docs/logs.
Reviewed changes
Copilot reviewed 8 out of 8 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
crates/ppvm-tableau/src/measure.rs |
Main performance refactor: mem::replace-based ownership transfer, case-b fast path (no HashMap), real-only overlap accumulation, retain-based A/B partition for case-a. |
crates/ppvm-tableau/src/data.rs |
Removes now-unused trim_coefficients_for_measurement() helper. |
crates/ppvm-tableau/benches/measure-scaling.rs |
Adds Criterion benchmark to track measurement scaling at different T-gate counts. |
crates/ppvm-tableau/Cargo.toml |
Registers the new measure-scaling benchmark target. |
crates/ppvm-tableau/examples/profile_measure.rs |
Adds an ad-hoc profiling example for end-to-end measurement time. |
docs/autotune/2026-04-12-measure-parallel/metric.toml |
Records perf results across micro-optimization iterations. |
docs/autotune/2026-04-12-measure-parallel/log.md |
Logs profiling breakdown and optimization notes. |
docs/autotune/2026-04-12-measure-parallel/eliminate-clone/prompts.md |
Captures the prompt/plan used for the “eliminate clone” iteration. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| let mut t = tab.fork(Some(42)); | ||
| let start = Instant::now(); | ||
| for q in 0..n_qubits { | ||
| t.measure(q); |
| || tab.fork(Some(42)), | ||
| |t| { | ||
| for q in 0..n_qubits { | ||
| t.measure(q); |
| // Partition into A (k-bit=0) and B (k-bit=1) via drain, then merge. | ||
| // This avoids allocating a separate b_keys Vec and eliminates | ||
| // individual HashMap removes (which require probe+shift). |
| | Component | 256 coeffs | 1024 coeffs | 4096 coeffs | 16384 coeffs | | ||
| |-----------|-----------|------------|------------|-------------| | ||
| | decomp | 1.8% | 0.6% | 0.1% | <0.1% | | ||
| | clone (coeffs→HashMap) | 23.9% | 26.8% | 36.1% | 26.3% | | ||
| | overlap loop | 21.1% | 22.7% | 17.6% | 22.4% | |
There was a problem hiding this comment.
Pull request overview
This PR optimizes GeneralizedTableau::measure() in ppvm-tableau, targeting the dominant end-to-end runtime cost for circuits with higher T-gate counts by reducing allocation/HashMap overhead and simplifying overlap accumulation.
Changes:
- Adds a case-b fast path (
stab_anticomm_bits == 0) that avoids building aHashMapand computes overlap/filtering directly from coefficient entries. - Refactors overlap computation to accumulate only the real component (and skips odd phases in the case-b fast path).
- Improves case-a drain/merge by partitioning with
HashMap::retaininstead of collecting/removing keys individually; adds a new Criterion bench and profiling/example + autotune docs.
Reviewed changes
Copilot reviewed 8 out of 8 changed files in this pull request and generated no comments.
Show a summary per file
| File | Description |
|---|---|
crates/ppvm-tableau/src/measure.rs |
Implements the measurement micro-optimizations (case-b fast path, real-only overlap, retain-based partitioning) and factors overlap/merge helpers. |
crates/ppvm-tableau/src/data.rs |
Exposes compute_phase_with_mask_static (and the rayon threshold) for reuse; removes now-unused measurement helper methods. |
crates/ppvm-tableau/benches/measure-scaling.rs |
Adds a Criterion benchmark to measure scaling across T-gate counts. |
crates/ppvm-tableau/Cargo.toml |
Registers the new measure-scaling benchmark target. |
crates/ppvm-tableau/examples/profile_measure.rs |
Adds an ad-hoc profiling example for measurement scaling. |
docs/autotune/2026-04-12-measure-parallel/* |
Captures autotuning metrics, logs, and prompts for the measurement optimization work. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Optimize GeneralizedTableau::measure(), which dominates end-to-end runtime at 85-90% for circuits with 8+ T gates, through five micro-optimizations: 1. Eliminate coefficient clone: replace clone().into_iter() with mem::replace() — self.coefficients is never read after the clone. 2. Case_b fast path: when Z is already a stabilizer (stab_anticomm_bits == 0), skip HashMap construction entirely. The overlap is self-pairing (conj(c)*c = |c|^2), so we work directly on the Vec with no HashMap needed. 3. Two-pass case_b: iterate by reference for overlap, then filter directly into coefficients — avoids the push-all-then-retain pattern. 4. Real-only overlap accumulation: compute only Re(z_overlap) since the imaginary part is always ~0. For case_b, skip odd-phase entries entirely (zero contribution to real part). 5. Retain-based drain: replace collect-b_keys + individual-remove with HashMap::retain to partition A/B entries in a single pass, avoiding expensive per-entry HashMap removes. Also: extract overlap and merge logic into helper methods, remove unused compute_phase_with_mask wrapper and trim_coefficients_for_measurement, make compute_phase_with_mask_static pub(crate). Benchmark results (85 qubits, full measurement of all qubits): | Benchmark | Before | After | Improvement | |------------------------|----------|----------|-------------| | MSD-fused (5 T gates) | 63.8 µs | 57.5 µs | -9.9% | | 8 T gates (256 coeffs) | 23.1 µs | 15.0 µs | -34.8% | | 10 T gates (1024) | 49.5 µs | 36.0 µs | -27.3% | | 12 T gates (4096) | 156.6 µs | 118.6 µs | -24.3% | Note: rayon parallelization was tested for measurement overlap and B→A merge but showed regressions at all sizes (16K-65K coefficients). The per-element work is too cheap (~5ns) for rayon overhead, and the sequential HashMap accumulation dominates. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
3801f4f to
6e4a193
Compare
Optimize GeneralizedTableau::measure(), which dominates end-to-end runtime at 85-90% for circuits with 8+ T gates, through five micro-optimizations: 1. Eliminate coefficient clone: replace clone().into_iter() with mem::replace() — self.coefficients is never read after the clone. 2. Case_b fast path: when Z is already a stabilizer (stab_anticomm_bits == 0), skip HashMap construction entirely. The overlap is self-pairing (conj(c)*c = |c|^2), so we work directly on the Vec with no HashMap needed. 3. Two-pass case_b: iterate by reference for overlap, then filter directly into coefficients — avoids the push-all-then-retain pattern. 4. Real-only overlap accumulation: compute only Re(z_overlap) since the imaginary part is always ~0. For case_b, skip odd-phase entries entirely (zero contribution to real part). 5. Retain-based drain: replace collect-b_keys + individual-remove with HashMap::retain to partition A/B entries in a single pass, avoiding expensive per-entry HashMap removes. Also: extract overlap and merge logic into helper methods, remove unused compute_phase_with_mask wrapper and trim_coefficients_for_measurement, make compute_phase_with_mask_static pub(crate). Benchmark results (85 qubits, full measurement of all qubits): | Benchmark | Before | After | Improvement | |------------------------|----------|----------|-------------| | MSD-fused (5 T gates) | 63.8 µs | 57.5 µs | -9.9% | | 8 T gates (256 coeffs) | 23.1 µs | 15.0 µs | -34.8% | | 10 T gates (1024) | 49.5 µs | 36.0 µs | -27.3% | | 12 T gates (4096) | 156.6 µs | 118.6 µs | -24.3% | Note: rayon parallelization was tested for measurement overlap and B→A merge but showed regressions at all sizes (16K-65K coefficients). The per-element work is too cheap (~5ns) for rayon overhead, and the sequential HashMap accumulation dominates. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Summary
Optimize the
GeneralizedTableau::measure()method which dominates end-to-end runtime at 85-90% for circuits with 8+ T gates. Five micro-optimizations that compose for 24-35% total improvement on measurement:clone().into_iter()withmem::replace()sinceself.coefficientsis never read after the clonestab_anticomm_bits == 0), skip HashMap construction entirely — the overlap is self-pairing (conj(c)*c = |c|^2)Re(z_overlap)since the imaginary part is always ~0. For case_b, skip odd-phase entries entirely (zero contribution)collect-b_keys + individual-removewithHashMap::retainto partition A/B entries in a single passBenchmark results (85 qubits, full measurement)
Test plan
cargo test -p ppvm-runtime -p ppvm-tableau -p ppvm-sym)test_t_gate_measurement_statisticsverifies measurement correctness🤖 Generated with Claude Code