Repository navigation
feat(tableau): add rayon feature for parallel coefficient branching - #55
Conversation
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds an optional `rayon` feature to ppvm-tableau that parallelizes the coefficient accumulation in `branch_with_coefficients` and `compute_coefficients_after_pauli_apply` using rayon's parallel fold/reduce. Key design decisions: - Feature-gated: all parallel code behind `features = ["rayon"]` - Threshold-based: only uses rayon when coefficient count >= 4096, below that sequential is always faster - Zero regression: when feature is off, code is identical to before (refactored to use extracted free functions but same logic) - Same API: no interface changes, just enable the feature Also adds Send + Sync + Copy + num::Num bounds to impl blocks across the crate (tgate, rot1, rot2, measure, stim, noise, reset) - all concrete types (f64, usize, u128, BUint) satisfy these trivially. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds a `large-tgate` benchmark group that isolates the cost of a single T-gate applied to pre-built states with 2K-131K coefficients. Also adds a profile_rayon example for quick A/B comparisons. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Replace the rayon fold/reduce-with-HashMap pattern (which was 2-3x slower due to thread-local HashMap merge overhead) with parallel map into a Vec of computed pairs followed by sequential accumulation into a pre-sized HashMap. Results at various coefficient counts (single T-gate application): - 32K coefficients: 35% faster than sequential - 131K coefficients: 39% faster - 1M coefficients: 47% faster Also raise threshold from 4096 to 16384 to avoid regressions at smaller coefficient counts where rayon overhead dominates. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
… 1M coeffs Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
Adds an optional rayon feature to ppvm-tableau to parallelize coefficient branching during non-Clifford operations, plus supporting benchmarks/profiling artifacts.
Changes:
- Introduces
rayonCargo feature/dependency and parallel coefficient-branching helpers indata.rs(parallel compute → sequential HashMap accumulate). - Tightens trait bounds across tableau operations to support parallel execution (
Send/Sync/Copy/num::Numadditions). - Adds new large-coefficient Criterion benchmarks and several profiling examples, plus autotune logs/metrics.
Reviewed changes
Copilot reviewed 16 out of 17 changed files in this pull request and generated 5 comments.
Show a summary per file
| File | Description |
|---|---|
| docs/autotune/2026-04-11-tableau-rayon-threading/metric.toml | Records baseline vs rayon performance metrics for coefficient branching experiments |
| docs/autotune/2026-04-11-tableau-rayon-threading/log.md | Design notes and profiling results motivating the chosen parallel strategy/threshold |
| crates/ppvm-tableau/Cargo.toml | Adds optional rayon dependency and rayon feature flag |
| Cargo.lock | Locks the new rayon dependency |
| crates/ppvm-tableau/src/data.rs | Implements parallelized branching/apply coefficient accumulation and refactors phase computation into a static helper |
| crates/ppvm-tableau/src/stim.rs | Updates bounds on RunStim impl to align with new tableau constraints |
| crates/ppvm-tableau/src/measure.rs | Updates bounds on measurement impl to align with new tableau constraints |
| crates/ppvm-tableau/src/noise.rs | Updates bounds on noise-channel impls to align with new tableau constraints |
| crates/ppvm-tableau/src/gates/tgate.rs | Updates bounds for T-gate impl to align with new tableau constraints |
| crates/ppvm-tableau/src/gates/rot1.rs | Updates bounds for rotation impl to align with new tableau constraints |
| crates/ppvm-tableau/src/gates/rot2.rs | Updates bounds for rotation impl to align with new tableau constraints |
| crates/ppvm-tableau/src/gates/reset.rs | Updates bounds for reset impl to align with new tableau constraints |
| crates/ppvm-tableau/benches/tableau.rs | Adds large-tgate benchmark group to isolate single-T branching at large coefficient counts |
| crates/ppvm-tableau/examples/profile_scaling.rs | Adds scaling profiler for Clifford gates, T branching, and measurement costs |
| crates/ppvm-tableau/examples/profile_rayon.rs | Adds sequential vs rayon profiling harness focused on T branching |
| crates/ppvm-tableau/examples/profile_msd.rs | Adds end-to-end MSD-like workload profiler |
| crates/ppvm-tableau/examples/profile_breakdown.rs | Adds micro-profiler to split compute vs accumulation time inside branching loop |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| // Only collect to Vec when rayon might be used; otherwise iterate directly | ||
| #[cfg(feature = "rayon")] | ||
| let new_coefficients = { | ||
| let items: Vec<_> = old_coefficients.into_iter().collect(); | ||
| branch_coefficients_parallel( |
There was a problem hiding this comment.
With feature = "rayon" enabled, this block always collects old_coefficients into a new Vec before deciding whether to use the parallel path. That means small coefficient counts (< threshold) still pay an allocation/move cost, so the threshold doesn’t actually avoid overhead. Consider branching on old_coefficients.len() first and calling the sequential accumulator directly when below the threshold, only collecting into a Vec for the parallel case.
| #[cfg(feature = "rayon")] | ||
| let new_coefficients = { | ||
| let items: Vec<_> = old_coefficients.into_iter().collect(); | ||
| apply_coefficients_parallel( | ||
| &items, |
There was a problem hiding this comment.
Same issue as branch_with_coefficients: under feature = "rayon" this always collects old_coefficients into a Vec, even when below the rayon threshold (where the sequential path is used). To keep the feature truly “no-regression” for small coefficient counts, branch on old_coefficients.len() and only allocate/collect when the parallel path will run.
| /// Accumulate branch coefficients. Uses rayon parallel fold/reduce when the feature | ||
| /// is enabled AND the coefficient count exceeds `RAYON_COEFF_THRESHOLD`. |
There was a problem hiding this comment.
The doc comment mentions a “rayon parallel fold/reduce” strategy, but the implementation now uses a parallel map/collect into a Vec followed by sequential accumulation. Updating the comment would avoid misleading future readers.
| /// Accumulate branch coefficients. Uses rayon parallel fold/reduce when the feature | |
| /// is enabled AND the coefficient count exceeds `RAYON_COEFF_THRESHOLD`. | |
| /// Accumulate branch coefficients. When rayon is enabled and the coefficient count | |
| /// exceeds `RAYON_COEFF_THRESHOLD`, this uses parallel mapping/collection into a | |
| /// `Vec` followed by sequential accumulation. |
| impl<T: Config, I, C: SparseVector<Complex<T::Coeff>, I>> GeneralizedTableau<T, I, C> | ||
| where | ||
| T::Coeff: One + Zero + Clone, | ||
| T::Coeff: One + Zero + Clone + Send + Sync + num::Num, | ||
| Complex<T::Coeff>: std::ops::Mul<Output = Complex<T::Coeff>> | ||
| + std::ops::AddAssign | ||
| + From<Complex64> | ||
| + ComplexFloat, | ||
| I: TableauIndex, | ||
| + ComplexFloat | ||
| + Copy, | ||
| I: TableauIndex + Send + Sync, |
There was a problem hiding this comment.
This impl now requires T::Coeff: num::Num (and Complex<T::Coeff>: Copy) unconditionally. Since Config::Coeff is only constrained by ppvm_runtime::traits::Coefficient, this tightens the public generic surface area even when the rayon feature is off. If these bounds are only needed for the parallel path, consider limiting them to the rayon-gated code (e.g., by adding method-level where bounds for the parallel helpers or using separate #[cfg(feature = "rayon")] helper impls) to avoid unnecessary API restrictions.
|
|
||
| large_group.bench_function(format!("single-t-on-{n_coeffs}-coeffs"), |b| { | ||
| b.iter_batched_ref( | ||
| || setup.fork(None), |
There was a problem hiding this comment.
The benchmark setup uses setup.fork(None), which reseeds the RNG from entropy each iteration. That adds extra overhead and variance unrelated to the T-gate branching cost being measured. Consider using a fixed seed (e.g. fork(Some(...))) or cloning without reseeding so the benchmark isolates the hot path more accurately.
| || setup.fork(None), | |
| || setup.fork(Some(0)), |
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
|
…x docs - Skip Vec::collect when coefficient count is below RAYON_COEFF_THRESHOLD, calling the sequential path directly (fixes unnecessary allocation) - Update doc comments to reflect parallel map/collect strategy (not fold/reduce) - Use deterministic seed in large-tgate benchmark (fork(Some(0)) not fork(None)) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Adds a `fused-tgate-circuit` benchmark group that combines batch Clifford operations (fusion) with variable T-gates and measurement. Also adds profile_combined example showing time breakdown. Key finding: measurement dominates at 85-90% of total circuit time when coefficient counts are large (>4K). Fusion reduces Clifford overhead to <1%. T-gate branching (rayon target) is ~10%. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 15 out of 16 changed files in this pull request and generated 4 comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| impl<T: Config, I, C: SparseVector<Complex<T::Coeff>, I>> GeneralizedTableau<T, I, C> | ||
| where | ||
| T::Coeff: One + Zero + Clone, | ||
| T::Coeff: One + Zero + Clone + Send + Sync + num::Num, | ||
| Complex<T::Coeff>: std::ops::Mul<Output = Complex<T::Coeff>> | ||
| + std::ops::AddAssign | ||
| + From<Complex64> | ||
| + ComplexFloat, | ||
| I: TableauIndex, | ||
| + ComplexFloat | ||
| + Copy, | ||
| I: TableauIndex + Send + Sync, | ||
| { |
There was a problem hiding this comment.
The impl GeneralizedTableau<...> block now adds extra bounds (T::Coeff: Send + Sync + num::Num, Complex<T::Coeff>: Copy, I: Send + Sync). Even if these are typically true, they change when inherent methods like new, fork, etc. are available, which contradicts the PR description’s “No API changes”. Consider keeping the minimal bounds on the main impl and putting the stronger Send+Sync requirements only on the #[cfg(feature = "rayon")] parallel helpers / code paths (e.g., via separate impl blocks or helper trait aliases).
| let pairs: Vec<(I, Complex<CoeffType>, I, Complex<CoeffType>)> = items | ||
| .par_iter() | ||
| .map(|&(coeff, idx)| { | ||
| let branch_index = idx ^ stab_anticomm_bits; | ||
| let branch_phase_contribution = compute_phase_with_mask_static( | ||
| destab_anticomm_bits, | ||
| idx, | ||
| stab_anticomm_bits, | ||
| odd_phase_mask, | ||
| ); | ||
| let branch_phase = (branch_phase_contribution + phase_decomp) % 4; | ||
| let phase_factor: Complex<CoeffType> = | ||
| COMPLEX_PHASE_CONVERSION[branch_phase as usize].into(); | ||
| ( | ||
| branch_index, | ||
| phase_factor * coeff * branch_factor, | ||
| idx, | ||
| coeff * coefficient_factor, | ||
| ) | ||
| }) |
There was a problem hiding this comment.
The parallel coefficient path drops the invariant check that exists in the sequential path (debug_assert! that coefficients are non-zero). To keep debug behavior consistent across feature flags, add the same debug_assert! inside the par_iter().map(...) closure (and similarly for the apply-coefficients parallel path).
| // When rayon is enabled and above the threshold, collect to Vec for parallel map; | ||
| // otherwise iterate directly with the sequential path. | ||
| #[cfg(feature = "rayon")] | ||
| let new_coefficients = if n_coefficients >= RAYON_COEFF_THRESHOLD { | ||
| let items: Vec<_> = old_coefficients.into_iter().collect(); | ||
| branch_coefficients_parallel( | ||
| &items, | ||
| stab_anticomm_bits, | ||
| destab_anticomm_bits, | ||
| idx, | ||
| odd_phase_mask, | ||
| phase_decomp, | ||
| coefficient_factor, | ||
| branch_factor, | ||
| ) | ||
| } else { | ||
| branch_coefficients_seq( | ||
| old_coefficients, | ||
| stab_anticomm_bits, | ||
| destab_anticomm_bits, | ||
| odd_phase_mask, | ||
| ); | ||
|
|
||
| // the total phase is the product of the above with the decomposition phase | ||
| let branch_phase = (branch_phase_contribution + phase_decomp) % 4; | ||
|
|
||
| let phase_factor: Complex<T::Coeff> = | ||
| COMPLEX_PHASE_CONVERSION[branch_phase as usize].into(); | ||
|
|
||
| let branch_coefficient = phase_factor * coeff * branch_factor; | ||
| let nonbranch_coefficient = coeff * coefficient_factor; | ||
| phase_decomp, | ||
| coefficient_factor, | ||
| branch_factor, | ||
| ) | ||
| }; |
There was a problem hiding this comment.
The new rayon branch is correctness-sensitive (parallel map + sequential accumulate) but there’s no unit test that exercises the items.len() >= RAYON_COEFF_THRESHOLD path and checks it matches the sequential result. Adding a #[cfg(feature = "rayon")] test in this module that compares *_parallel vs *_seq outputs for a deterministic synthetic input (size just over the threshold) would help prevent future regressions.
| let pairs: Vec<(I, Complex<CoeffType>)> = items | ||
| .par_iter() | ||
| .map(|&(coeff, idx)| { | ||
| let branch_index = idx ^ stab_anticomm_bits; | ||
| let branch_phase_contribution = compute_phase_with_mask_static( | ||
| destab_anticomm_bits, | ||
| idx, | ||
| stab_anticomm_bits, | ||
| odd_phase_mask, | ||
| ); | ||
| let branch_phase = (branch_phase_contribution + phase_decomp) % 4; | ||
| let phase_factor: Complex<CoeffType> = | ||
| COMPLEX_PHASE_CONVERSION[branch_phase as usize].into(); | ||
| (branch_index, phase_factor * coeff) | ||
| }) | ||
| .collect(); |
There was a problem hiding this comment.
Like branch_coefficients_parallel, this apply_coefficients_parallel path also lacks the sequential-path debug_assert! that coefficients are non-zero. Add the same assertion inside the parallel mapping closure to keep debug invariants consistent when the rayon feature is enabled.
Summary
rayoncargo feature toppvm-tableauthat parallelizes coefficient branching in non-Clifford gate application (branch_with_coefficients,compute_coefficients_after_pauli_apply)(key, value)pairs in parallel, then a sequential loop inserts into a pre-sized HashMaplarge-tgategroup) and profiling examplesPerformance (single T-gate on pre-built state)
rayonUsage
No API changes — just enable the feature and parallel execution kicks in automatically for large coefficient sets.
Test plan
cargo test -p ppvm-tableau --lib— 117 tests pass (no feature)cargo test -p ppvm-tableau --lib --features rayon— 117 tests pass (with feature)cargo check --workspace— full workspace compilescargo bench -p ppvm-tableau --bench tableau -- "tableau-t-gate"— no regression at small countscargo bench -p ppvm-tableau --bench tableau --features rayon -- "large-tgate"— speedup at large counts🤖 Generated with Claude Code