Skip to content

perf(tableau): optimize measurement for generalized tableau - #56

Merged
Roger-luo merged 2 commits into
mainfrom
autotune/measure-parallel
Apr 13, 2026
Merged

Roger-luo merged 2 commits into
mainfrom
autotune/measure-parallel

Conversation

@Roger-luo

Copy link
Copy Markdown
Collaborator

Summary

Optimize the GeneralizedTableau::measure() method which dominates end-to-end runtime at 85-90% for circuits with 8+ T gates. Five micro-optimizations that compose for 24-35% total improvement on measurement:

  • Eliminate coefficient clone: Replace clone().into_iter() with mem::replace() since self.coefficients is never read after the clone
  • Case_b fast path: When Z is already a stabilizer (stab_anticomm_bits == 0), skip HashMap construction entirely — the overlap is self-pairing (conj(c)*c = |c|^2)
  • Avoid retain: Two-pass approach for case_b — iterate by reference for overlap, then filter directly into coefficients
  • Real-only accumulation: Compute only Re(z_overlap) since the imaginary part is always ~0. For case_b, skip odd-phase entries entirely (zero contribution)
  • Retain-based drain: Replace collect-b_keys + individual-remove with HashMap::retain to partition A/B entries in a single pass

Benchmark results (85 qubits, full measurement)

Benchmark Before After Improvement
MSD-fused (5 T gates) 63.8µs 57.5µs -9.9%
8 T gates (256 coeffs) 23.1µs 15.0µs -34.8%
10 T gates (1024 coeffs) 49.5µs 36.0µs -27.3%
12 T gates (4096 coeffs) 156.6µs 118.6µs -24.3%

Test plan

  • All 372 Rust tests pass (cargo test -p ppvm-runtime -p ppvm-tableau -p ppvm-sym)
  • Criterion benchmarks show consistent improvement across T-gate counts
  • test_t_gate_measurement_statistics verifies measurement correctness
  • Review measurement logic for case_b self-pairing simplification

🤖 Generated with Claude Code

Copilot AI review requested due to automatic review settings April 12, 2026 23:44
@github-actions

github-actions Bot commented Apr 12, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-04-13 01:22 UTC

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR optimizes GeneralizedTableau::measure() (the dominant runtime hotspot for generalized tableau simulations with multiple T gates) by reducing allocations and HashMap work, and by simplifying overlap computation to accumulate only the real component.

Changes:

  • Reworks GeneralizedTableau::measure() to avoid cloning coefficients, add a case-b fast path, and optimize case-a partitioning/drain behavior.
  • Removes trim_coefficients_for_measurement() (now inlined/handled in measure()’s case-b path).
  • Adds a new Criterion benchmark (measure-scaling) and an ad-hoc profiling example, plus autotune docs/logs.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 4 comments.

Show a summary per file
File Description
crates/ppvm-tableau/src/measure.rs Main performance refactor: mem::replace-based ownership transfer, case-b fast path (no HashMap), real-only overlap accumulation, retain-based A/B partition for case-a.
crates/ppvm-tableau/src/data.rs Removes now-unused trim_coefficients_for_measurement() helper.
crates/ppvm-tableau/benches/measure-scaling.rs Adds Criterion benchmark to track measurement scaling at different T-gate counts.
crates/ppvm-tableau/Cargo.toml Registers the new measure-scaling benchmark target.
crates/ppvm-tableau/examples/profile_measure.rs Adds an ad-hoc profiling example for end-to-end measurement time.
docs/autotune/2026-04-12-measure-parallel/metric.toml Records perf results across micro-optimization iterations.
docs/autotune/2026-04-12-measure-parallel/log.md Logs profiling breakdown and optimization notes.
docs/autotune/2026-04-12-measure-parallel/eliminate-clone/prompts.md Captures the prompt/plan used for the “eliminate clone” iteration.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

let mut t = tab.fork(Some(42));
let start = Instant::now();
for q in 0..n_qubits {
t.measure(q);
|| tab.fork(Some(42)),
|t| {
for q in 0..n_qubits {
t.measure(q);
Comment thread crates/ppvm-tableau/src/measure.rs Outdated
Comment on lines +194 to +196
// Partition into A (k-bit=0) and B (k-bit=1) via drain, then merge.
// This avoids allocating a separate b_keys Vec and eliminates
// individual HashMap removes (which require probe+shift).
Comment on lines +9 to +13
| Component | 256 coeffs | 1024 coeffs | 4096 coeffs | 16384 coeffs |
|-----------|-----------|------------|------------|-------------|
| decomp | 1.8% | 0.6% | 0.1% | <0.1% |
| clone (coeffs→HashMap) | 23.9% | 26.8% | 36.1% | 26.3% |
| overlap loop | 21.1% | 22.7% | 17.6% | 22.4% |
Copilot AI review requested due to automatic review settings April 13, 2026 01:12

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR optimizes GeneralizedTableau::measure() in ppvm-tableau, targeting the dominant end-to-end runtime cost for circuits with higher T-gate counts by reducing allocation/HashMap overhead and simplifying overlap accumulation.

Changes:

  • Adds a case-b fast path (stab_anticomm_bits == 0) that avoids building a HashMap and computes overlap/filtering directly from coefficient entries.
  • Refactors overlap computation to accumulate only the real component (and skips odd phases in the case-b fast path).
  • Improves case-a drain/merge by partitioning with HashMap::retain instead of collecting/removing keys individually; adds a new Criterion bench and profiling/example + autotune docs.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated no comments.

Show a summary per file
File Description
crates/ppvm-tableau/src/measure.rs Implements the measurement micro-optimizations (case-b fast path, real-only overlap, retain-based partitioning) and factors overlap/merge helpers.
crates/ppvm-tableau/src/data.rs Exposes compute_phase_with_mask_static (and the rayon threshold) for reuse; removes now-unused measurement helper methods.
crates/ppvm-tableau/benches/measure-scaling.rs Adds a Criterion benchmark to measure scaling across T-gate counts.
crates/ppvm-tableau/Cargo.toml Registers the new measure-scaling benchmark target.
crates/ppvm-tableau/examples/profile_measure.rs Adds an ad-hoc profiling example for measurement scaling.
docs/autotune/2026-04-12-measure-parallel/* Captures autotuning metrics, logs, and prompts for the measurement optimization work.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Roger-luo and others added 2 commits April 12, 2026 21:18
Optimize GeneralizedTableau::measure(), which dominates end-to-end
runtime at 85-90% for circuits with 8+ T gates, through five
micro-optimizations:

1. Eliminate coefficient clone: replace clone().into_iter() with
   mem::replace() — self.coefficients is never read after the clone.

2. Case_b fast path: when Z is already a stabilizer
   (stab_anticomm_bits == 0), skip HashMap construction entirely.
   The overlap is self-pairing (conj(c)*c = |c|^2), so we work
   directly on the Vec with no HashMap needed.

3. Two-pass case_b: iterate by reference for overlap, then filter
   directly into coefficients — avoids the push-all-then-retain
   pattern.

4. Real-only overlap accumulation: compute only Re(z_overlap) since
   the imaginary part is always ~0. For case_b, skip odd-phase
   entries entirely (zero contribution to real part).

5. Retain-based drain: replace collect-b_keys + individual-remove
   with HashMap::retain to partition A/B entries in a single pass,
   avoiding expensive per-entry HashMap removes.

Also: extract overlap and merge logic into helper methods, remove
unused compute_phase_with_mask wrapper and trim_coefficients_for_measurement,
make compute_phase_with_mask_static pub(crate).

Benchmark results (85 qubits, full measurement of all qubits):

  | Benchmark              | Before   | After    | Improvement |
  |------------------------|----------|----------|-------------|
  | MSD-fused (5 T gates)  | 63.8 µs  | 57.5 µs  | -9.9%       |
  | 8 T gates (256 coeffs) | 23.1 µs  | 15.0 µs  | -34.8%      |
  | 10 T gates (1024)      | 49.5 µs  | 36.0 µs  | -27.3%      |
  | 12 T gates (4096)      | 156.6 µs | 118.6 µs | -24.3%      |

Note: rayon parallelization was tested for measurement overlap and
B→A merge but showed regressions at all sizes (16K-65K coefficients).
The per-element work is too cheap (~5ns) for rayon overhead, and the
sequential HashMap accumulation dominates.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@Roger-luo
Roger-luo force-pushed the autotune/measure-parallel branch from 3801f4f to 6e4a193 Compare April 13, 2026 01:19
@Roger-luo
Roger-luo merged commit 03437f6 into main Apr 13, 2026
7 checks passed
Roger-luo added a commit that referenced this pull request Apr 13, 2026
Optimize GeneralizedTableau::measure(), which dominates end-to-end
runtime at 85-90% for circuits with 8+ T gates, through five
micro-optimizations:

1. Eliminate coefficient clone: replace clone().into_iter() with
   mem::replace() — self.coefficients is never read after the clone.

2. Case_b fast path: when Z is already a stabilizer
   (stab_anticomm_bits == 0), skip HashMap construction entirely.
   The overlap is self-pairing (conj(c)*c = |c|^2), so we work
   directly on the Vec with no HashMap needed.

3. Two-pass case_b: iterate by reference for overlap, then filter
   directly into coefficients — avoids the push-all-then-retain
   pattern.

4. Real-only overlap accumulation: compute only Re(z_overlap) since
   the imaginary part is always ~0. For case_b, skip odd-phase
   entries entirely (zero contribution to real part).

5. Retain-based drain: replace collect-b_keys + individual-remove
   with HashMap::retain to partition A/B entries in a single pass,
   avoiding expensive per-entry HashMap removes.

Also: extract overlap and merge logic into helper methods, remove
unused compute_phase_with_mask wrapper and trim_coefficients_for_measurement,
make compute_phase_with_mask_static pub(crate).

Benchmark results (85 qubits, full measurement of all qubits):

  | Benchmark              | Before   | After    | Improvement |
  |------------------------|----------|----------|-------------|
  | MSD-fused (5 T gates)  | 63.8 µs  | 57.5 µs  | -9.9%       |
  | 8 T gates (256 coeffs) | 23.1 µs  | 15.0 µs  | -34.8%      |
  | 10 T gates (1024)      | 49.5 µs  | 36.0 µs  | -27.3%      |
  | 12 T gates (4096)      | 156.6 µs | 118.6 µs | -24.3%      |

Note: rayon parallelization was tested for measurement overlap and
B→A merge but showed regressions at all sizes (16K-65K coefficients).
The per-element work is too cheap (~5ns) for rayon overhead, and the
sequential HashMap accumulation dominates.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@Roger-luo
Roger-luo deleted the autotune/measure-parallel branch April 13, 2026 01:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants