Skip to content

feat(tableau): add rayon feature for parallel coefficient branching - #55

Merged
Roger-luo merged 9 commits into
mainfrom
autotune/tableau-rayon-threading
Apr 12, 2026
Merged

Roger-luo merged 9 commits into
mainfrom
autotune/tableau-rayon-threading

Conversation

@Roger-luo

Copy link
Copy Markdown
Collaborator

Summary

  • Adds an optional rayon cargo feature to ppvm-tableau that parallelizes coefficient branching in non-Clifford gate application (branch_with_coefficients, compute_coefficients_after_pauli_apply)
  • Uses parallel map → Vec + sequential accumulate pattern: rayon computes all (key, value) pairs in parallel, then a sequential loop inserts into a pre-sized HashMap
  • Threshold-based: only engages rayon when coefficient count ≥ 16,384 — below that, the sequential path is always used
  • Zero regression when feature is off: all 117 tests pass, benchmarks identical to baseline
  • Adds large-coefficient T-gate benchmarks (large-tgate group) and profiling examples

Performance (single T-gate on pre-built state)

Coefficients Sequential With rayon Speedup
≤8,192 — — No change (below threshold)
32,768 1.15 ms 748 µs +35%
131,072 4.14 ms 2.52 ms +39%
1,048,576 45.8 ms 24.1 ms +47%

Usage

[dependencies]
ppvm-tableau = { version = "0.1.0", features = ["rayon"] }

No API changes — just enable the feature and parallel execution kicks in automatically for large coefficient sets.

Test plan

  • cargo test -p ppvm-tableau --lib — 117 tests pass (no feature)
  • cargo test -p ppvm-tableau --lib --features rayon — 117 tests pass (with feature)
  • cargo check --workspace — full workspace compiles
  • cargo bench -p ppvm-tableau --bench tableau -- "tableau-t-gate" — no regression at small counts
  • cargo bench -p ppvm-tableau --bench tableau --features rayon -- "large-tgate" — speedup at large counts

🤖 Generated with Claude Code

Roger-luo and others added 6 commits April 11, 2026 23:52
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds an optional `rayon` feature to ppvm-tableau that parallelizes the
coefficient accumulation in `branch_with_coefficients` and
`compute_coefficients_after_pauli_apply` using rayon's parallel fold/reduce.

Key design decisions:
- Feature-gated: all parallel code behind `features = ["rayon"]`
- Threshold-based: only uses rayon when coefficient count >= 4096,
  below that sequential is always faster
- Zero regression: when feature is off, code is identical to before
  (refactored to use extracted free functions but same logic)
- Same API: no interface changes, just enable the feature

Also adds Send + Sync + Copy + num::Num bounds to impl blocks across
the crate (tgate, rot1, rot2, measure, stim, noise, reset) - all
concrete types (f64, usize, u128, BUint) satisfy these trivially.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds a `large-tgate` benchmark group that isolates the cost of a single
T-gate applied to pre-built states with 2K-131K coefficients. Also adds
a profile_rayon example for quick A/B comparisons.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace the rayon fold/reduce-with-HashMap pattern (which was 2-3x
slower due to thread-local HashMap merge overhead) with parallel map
into a Vec of computed pairs followed by sequential accumulation into
a pre-sized HashMap.

Results at various coefficient counts (single T-gate application):
- 32K coefficients: 35% faster than sequential
- 131K coefficients: 39% faster
- 1M coefficients: 47% faster

Also raise threshold from 4096 to 16384 to avoid regressions at
smaller coefficient counts where rayon overhead dominates.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
… 1M coeffs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings April 12, 2026 19:04

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds an optional rayon feature to ppvm-tableau to parallelize coefficient branching during non-Clifford operations, plus supporting benchmarks/profiling artifacts.

Changes:

  • Introduces rayon Cargo feature/dependency and parallel coefficient-branching helpers in data.rs (parallel compute → sequential HashMap accumulate).
  • Tightens trait bounds across tableau operations to support parallel execution (Send/Sync/Copy/num::Num additions).
  • Adds new large-coefficient Criterion benchmarks and several profiling examples, plus autotune logs/metrics.

Reviewed changes

Copilot reviewed 16 out of 17 changed files in this pull request and generated 5 comments.

Show a summary per file
File Description
docs/autotune/2026-04-11-tableau-rayon-threading/metric.toml Records baseline vs rayon performance metrics for coefficient branching experiments
docs/autotune/2026-04-11-tableau-rayon-threading/log.md Design notes and profiling results motivating the chosen parallel strategy/threshold
crates/ppvm-tableau/Cargo.toml Adds optional rayon dependency and rayon feature flag
Cargo.lock Locks the new rayon dependency
crates/ppvm-tableau/src/data.rs Implements parallelized branching/apply coefficient accumulation and refactors phase computation into a static helper
crates/ppvm-tableau/src/stim.rs Updates bounds on RunStim impl to align with new tableau constraints
crates/ppvm-tableau/src/measure.rs Updates bounds on measurement impl to align with new tableau constraints
crates/ppvm-tableau/src/noise.rs Updates bounds on noise-channel impls to align with new tableau constraints
crates/ppvm-tableau/src/gates/tgate.rs Updates bounds for T-gate impl to align with new tableau constraints
crates/ppvm-tableau/src/gates/rot1.rs Updates bounds for rotation impl to align with new tableau constraints
crates/ppvm-tableau/src/gates/rot2.rs Updates bounds for rotation impl to align with new tableau constraints
crates/ppvm-tableau/src/gates/reset.rs Updates bounds for reset impl to align with new tableau constraints
crates/ppvm-tableau/benches/tableau.rs Adds large-tgate benchmark group to isolate single-T branching at large coefficient counts
crates/ppvm-tableau/examples/profile_scaling.rs Adds scaling profiler for Clifford gates, T branching, and measurement costs
crates/ppvm-tableau/examples/profile_rayon.rs Adds sequential vs rayon profiling harness focused on T branching
crates/ppvm-tableau/examples/profile_msd.rs Adds end-to-end MSD-like workload profiler
crates/ppvm-tableau/examples/profile_breakdown.rs Adds micro-profiler to split compute vs accumulation time inside branching loop

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread crates/ppvm-tableau/src/data.rs Outdated
Comment on lines +811 to +815
// Only collect to Vec when rayon might be used; otherwise iterate directly
#[cfg(feature = "rayon")]
let new_coefficients = {
let items: Vec<_> = old_coefficients.into_iter().collect();
branch_coefficients_parallel(

Copilot AI Apr 12, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With feature = "rayon" enabled, this block always collects old_coefficients into a new Vec before deciding whether to use the parallel path. That means small coefficient counts (< threshold) still pay an allocation/move cost, so the threshold doesn’t actually avoid overhead. Consider branching on old_coefficients.len() first and calling the sequential accumulator directly when below the threshold, only collecting into a Vec for the parallel case.

Copilot uses AI. Check for mistakes.
Comment on lines +867 to +871
#[cfg(feature = "rayon")]
let new_coefficients = {
let items: Vec<_> = old_coefficients.into_iter().collect();
apply_coefficients_parallel(
&items,

Copilot AI Apr 12, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same issue as branch_with_coefficients: under feature = "rayon" this always collects old_coefficients into a Vec, even when below the rayon threshold (where the sequential path is used). To keep the feature truly “no-regression” for small coefficient counts, branch on old_coefficients.len() and only allocate/collect when the parallel path will run.

Copilot uses AI. Check for mistakes.
Comment thread crates/ppvm-tableau/src/data.rs Outdated
Comment on lines +344 to +345
/// Accumulate branch coefficients. Uses rayon parallel fold/reduce when the feature
/// is enabled AND the coefficient count exceeds `RAYON_COEFF_THRESHOLD`.

Copilot AI Apr 12, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The doc comment mentions a “rayon parallel fold/reduce” strategy, but the implementation now uses a parallel map/collect into a Vec followed by sequential accumulation. Updating the comment would avoid misleading future readers.

Suggested change
/// Accumulate branch coefficients. Uses rayon parallel fold/reduce when the feature
/// is enabled AND the coefficient count exceeds `RAYON_COEFF_THRESHOLD`.
/// Accumulate branch coefficients. When rayon is enabled and the coefficient count
/// exceeds `RAYON_COEFF_THRESHOLD`, this uses parallel mapping/collection into a
/// `Vec` followed by sequential accumulation.

Copilot uses AI. Check for mistakes.
Comment on lines 514 to +522
impl<T: Config, I, C: SparseVector<Complex<T::Coeff>, I>> GeneralizedTableau<T, I, C>
where
T::Coeff: One + Zero + Clone,
T::Coeff: One + Zero + Clone + Send + Sync + num::Num,
Complex<T::Coeff>: std::ops::Mul<Output = Complex<T::Coeff>>
+ std::ops::AddAssign
+ From<Complex64>
+ ComplexFloat,
I: TableauIndex,
+ ComplexFloat
+ Copy,
I: TableauIndex + Send + Sync,

Copilot AI Apr 12, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This impl now requires T::Coeff: num::Num (and Complex<T::Coeff>: Copy) unconditionally. Since Config::Coeff is only constrained by ppvm_runtime::traits::Coefficient, this tightens the public generic surface area even when the rayon feature is off. If these bounds are only needed for the parallel path, consider limiting them to the rayon-gated code (e.g., by adding method-level where bounds for the parallel helpers or using separate #[cfg(feature = "rayon")] helper impls) to avoid unnecessary API restrictions.

Copilot uses AI. Check for mistakes.
Comment thread crates/ppvm-tableau/benches/tableau.rs Outdated

large_group.bench_function(format!("single-t-on-{n_coeffs}-coeffs"), |b| {
b.iter_batched_ref(
|| setup.fork(None),

Copilot AI Apr 12, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The benchmark setup uses setup.fork(None), which reseeds the RNG from entropy each iteration. That adds extra overhead and variance unrelated to the T-gate branching cost being measured. Consider using a fixed seed (e.g. fork(Some(...))) or cloning without reseeding so the benchmark isolates the hot path more accurately.

Suggested change
|| setup.fork(None),
|| setup.fork(Some(0)),

Copilot uses AI. Check for mistakes.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@github-actions

github-actions Bot commented Apr 12, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-04-12 23:08 UTC

…x docs

- Skip Vec::collect when coefficient count is below RAYON_COEFF_THRESHOLD,
  calling the sequential path directly (fixes unnecessary allocation)
- Update doc comments to reflect parallel map/collect strategy (not fold/reduce)
- Use deterministic seed in large-tgate benchmark (fork(Some(0)) not fork(None))

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings April 12, 2026 23:00
Adds a `fused-tgate-circuit` benchmark group that combines batch
Clifford operations (fusion) with variable T-gates and measurement.
Also adds profile_combined example showing time breakdown.

Key finding: measurement dominates at 85-90% of total circuit time
when coefficient counts are large (>4K). Fusion reduces Clifford
overhead to <1%. T-gate branching (rayon target) is ~10%.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
@Roger-luo
Roger-luo enabled auto-merge (squash) April 12, 2026 23:06

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 15 out of 16 changed files in this pull request and generated 4 comments.


💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines 516 to 525
impl<T: Config, I, C: SparseVector<Complex<T::Coeff>, I>> GeneralizedTableau<T, I, C>
where
T::Coeff: One + Zero + Clone,
T::Coeff: One + Zero + Clone + Send + Sync + num::Num,
Complex<T::Coeff>: std::ops::Mul<Output = Complex<T::Coeff>>
+ std::ops::AddAssign
+ From<Complex64>
+ ComplexFloat,
I: TableauIndex,
+ ComplexFloat
+ Copy,
I: TableauIndex + Send + Sync,
{

Copilot AI Apr 12, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The impl GeneralizedTableau<...> block now adds extra bounds (T::Coeff: Send + Sync + num::Num, Complex<T::Coeff>: Copy, I: Send + Sync). Even if these are typically true, they change when inherent methods like new, fork, etc. are available, which contradicts the PR description’s “No API changes”. Consider keeping the minimal bounds on the main impl and putting the stronger Send+Sync requirements only on the #[cfg(feature = "rayon")] parallel helpers / code paths (e.g., via separate impl blocks or helper trait aliases).

Copilot uses AI. Check for mistakes.
Comment on lines +368 to +387
let pairs: Vec<(I, Complex<CoeffType>, I, Complex<CoeffType>)> = items
.par_iter()
.map(|&(coeff, idx)| {
let branch_index = idx ^ stab_anticomm_bits;
let branch_phase_contribution = compute_phase_with_mask_static(
destab_anticomm_bits,
idx,
stab_anticomm_bits,
odd_phase_mask,
);
let branch_phase = (branch_phase_contribution + phase_decomp) % 4;
let phase_factor: Complex<CoeffType> =
COMPLEX_PHASE_CONVERSION[branch_phase as usize].into();
(
branch_index,
phase_factor * coeff * branch_factor,
idx,
coeff * coefficient_factor,
)
})

Copilot AI Apr 12, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The parallel coefficient path drops the invariant check that exists in the sequential path (debug_assert! that coefficients are non-zero). To keep debug behavior consistent across feature flags, add the same debug_assert! inside the par_iter().map(...) closure (and similarly for the apply-coefficients parallel path).

Copilot uses AI. Check for mistakes.
Comment on lines +815 to +839
// When rayon is enabled and above the threshold, collect to Vec for parallel map;
// otherwise iterate directly with the sequential path.
#[cfg(feature = "rayon")]
let new_coefficients = if n_coefficients >= RAYON_COEFF_THRESHOLD {
let items: Vec<_> = old_coefficients.into_iter().collect();
branch_coefficients_parallel(
&items,
stab_anticomm_bits,
destab_anticomm_bits,
idx,
odd_phase_mask,
phase_decomp,
coefficient_factor,
branch_factor,
)
} else {
branch_coefficients_seq(
old_coefficients,
stab_anticomm_bits,
destab_anticomm_bits,
odd_phase_mask,
);

// the total phase is the product of the above with the decomposition phase
let branch_phase = (branch_phase_contribution + phase_decomp) % 4;

let phase_factor: Complex<T::Coeff> =
COMPLEX_PHASE_CONVERSION[branch_phase as usize].into();

let branch_coefficient = phase_factor * coeff * branch_factor;
let nonbranch_coefficient = coeff * coefficient_factor;
phase_decomp,
coefficient_factor,
branch_factor,
)
};

Copilot AI Apr 12, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The new rayon branch is correctness-sensitive (parallel map + sequential accumulate) but there’s no unit test that exercises the items.len() >= RAYON_COEFF_THRESHOLD path and checks it matches the sequential result. Adding a #[cfg(feature = "rayon")] test in this module that compares *_parallel vs *_seq outputs for a deterministic synthetic input (size just over the threshold) would help prevent future regressions.

Copilot uses AI. Check for mistakes.
Comment on lines +468 to +483
let pairs: Vec<(I, Complex<CoeffType>)> = items
.par_iter()
.map(|&(coeff, idx)| {
let branch_index = idx ^ stab_anticomm_bits;
let branch_phase_contribution = compute_phase_with_mask_static(
destab_anticomm_bits,
idx,
stab_anticomm_bits,
odd_phase_mask,
);
let branch_phase = (branch_phase_contribution + phase_decomp) % 4;
let phase_factor: Complex<CoeffType> =
COMPLEX_PHASE_CONVERSION[branch_phase as usize].into();
(branch_index, phase_factor * coeff)
})
.collect();

Copilot AI Apr 12, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Like branch_coefficients_parallel, this apply_coefficients_parallel path also lacks the sequential-path debug_assert! that coefficients are non-zero. Add the same assertion inside the parallel mapping closure to keep debug invariants consistent when the rayon feature is enabled.

Copilot uses AI. Check for mistakes.
@Roger-luo
Roger-luo merged commit 8b0b963 into main Apr 12, 2026
7 checks passed
@Roger-luo
Roger-luo deleted the autotune/tableau-rayon-threading branch April 12, 2026 23:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants