Shrink affine addition scratch and improve prepared-point prefetching - #518
Merged
Merged
Conversation
|
Note Complete: Audit complete. V12 did not find any issues that need review. Open the full results here. Analyzed three files, diff |
ValarDragon
force-pushed
the
perf/affine-pending-prefetch-20260928
branch
from
September 27, 2026 23:28
f2947c0 to
3ce82f4
Compare
lovesh
added a commit
to PolymeshAssociation/arkworks-algebra
that referenced
this pull request
Sep 28, 2026
Batch-affine Pippenger finishes each level's additions in two passes: the prefix pass scales every numerator by its lane's product of earlier denominators, and the backward pass reads the slope directly and completes the addition. This replaces the separate batch inversion, its per-level prefix-product allocation, the pass that multiplied inverses by numerators, and the serial zero-denominator scan. After Zakura arkworks-rs#518 (zakura-core/common#518). msm_batch_affine on Pallas is 4-6% faster natively from 2^8 bases and 1-3% faster on wasm32. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ValarDragon
pushed a commit
that referenced
this pull request
Sep 28, 2026
Bump the workspace version and intra-workspace requirements from 2.0.0 to 2.1.0-rc.0 and assemble the v2.1.0-rc.0 changelog sections, consuming all 29 pending fragments. This is the first release of zakura-equihash, imported in #44. Editorial changes to the assembled entries: - Consolidate the ten incremental zakura-equihash solver speedup entries (#502–#516) into one entry, and give the C-compiler removal from #509 its own line: the crate ships for the first time, so its section describes the delta from upstream equihash 0.3.0 rather than the development history. - Merge the three zakura-bls12-381 final-exponentiation entries (#485, #494, #501) into one that states the variable-time inversion and the public-input timing contract, and drop the bare "#494" cross-reference. - Remove technique-only wording from the #483 (bellman), #489 (bls12_381), and #518 (pasta_curves) entries; #518 claims no speedup, so only the reduced scratch memory remains. - Rephrase the #494 bellman entry in the past tense used elsewhere. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Extracted from #507 into an independent PR against
main. This contains noIFMA arithmetic, VK evaluation-plan, IPA coefficient, or decoder changes.
64-bit Pasta targets. Prescale numerators in the prefix pass to remove a
multiplication from the dependent backward path without changing the total
multiplication count.
16 records ahead, with bounded dense, sparse, and range indexing.
no_stddetector, so this split does not depend on the IFMA PR. The prefetchinstruction itself is baseline-safe on x86-64.
No proof, transcript, protocol, or feature changes. Scalar exceptional-case
fallback and failure atomicity are preserved.
Validation
Local ARM, Rust 1.91.0:
glv,orbitsandmulticoreconfigurations checked.git diff --check, and independent arithmetic, pointer-bound,dependency, and API review passed.
Tests exercise non-unit prefixes, late zero denominators, odd lanes, every
point-preservation guard, dense/sparse/range staging, and detector bit boundaries.
The existing unrelated
SqrtTableHelpersdead-code warning remains in theno-default-feature configuration.
Linux native, Rust 1.97.1, rebased onto the immediate pre-#518 main:
prefetch_detection_matches_standard_libraryreports prefetch enabled.glv,orbitscheck passed, with the existingunrelated
SqrtTableHelperswarning.and rejected altered B1/B2 instances and canonical IPA blinding scalars.
The Pasta prepared record stays 96 bytes, so its table footprint does not grow.
Isolated prover and verifier impact
Compare main
f2c0f322with only this PR applied. The candidate's productiontree is byte-identical to merged #518 (
0c771246). No IFMA, IPA/VK, or decoderchanges are included. Both use the same diagnostic-only
benchmark.rsoverlay (SHA-256
fa020e7f70f18c7a163e26534427e30b05d6efc57b098fe089af6e149cbbf224).linux-1, AMD EPYC 9654 VMware guest, pinned CPU 2, one Rayon worker,Rust 1.97.1 release, default features plus
circuit,x86_64-asm,orbits.Builds used separate fresh target directories and fixed, hashed executables.
The initial shared-target build was rejected before any measurements because
Cargo reused the control binary for the candidate.
Proving uses one A/B/B/A bracket, ten Criterion samples per workload per
process, two-second warmup and 15-second requested measurement. Values below
are pooled medians of
times / iters, twenty samples per arm. Key generation,preparation and fixture construction are outside timing. Prepared proving is
warmed steady-state, not first-proof latency. The prepared and unprepared
harnesses differ in RNG setup; compare each mode against its own control.
Verification uses two complete A/B/B/A brackets, three warmups and fifteen
samples per size per process, using the same immutable corpus. Each bracket
has thirty samples per arm. Preparation, corpus loading, entry cloning,
positive validation and rejection checks are outside timing. Each proof has
two real Actions, so batch size two means four Actions.
All raw samples are retained. The first prepared single-proof control had a
seven-sample scheduling excursion; this motivated the complete verifier
repeat, not selective sample removal. The prepared 64-proof difference is
roughly neutral given VM noise; sub-1% improvements are small observations,
not significance claims. These results do not establish default-feature,
other-hardware, first-proof, or multicore performance.
The two initial control qualifications preceded an interruption by roughly
eight hours. The resumed comparison had fresh idle-host telemetry and fresh
balanced controls; qualification is not claimed to be immediately adjacent.
Frozen diagnostic revisions: control
0eaf9fa4, candidate0785578b.Corpus SHA-256:
c8bd4e97159e14cb1bafef13c482c82cc1dcac23efe7cf03d3e8cab3513cfd8e.Raw data and binary hashes are retained under:
/home/valar/benchmarks/pr518-affine-prover-verifier-20260928/logs/linux-1/home/valar/benchmarks/pr518-verifier-repeat-20260928/logs/linux-1The former combined #507 2x benchmark is not attributed to this PR.
API surface
No function, signature, visibility, trait, enum, or feature additions.
The existing
pub(crate) PreparedPoint<F>gains#[repr(align(32))];Pasta alignment changes from 8 to 32 bytes while the record remains 96 bytes.
Smaller generic-field records can gain padding. The prefetch helpers and
CPU/OS detector remain private.
The corresponding changes were removed from #507. The separate IPA/VK
split is #517; #507 retains its IFMA backend and existing decoder.