Skip to content

Shrink affine addition scratch and improve prepared-point prefetching - #518

Merged
ValarDragon merged 1 commit into
mainfrom
perf/affine-pending-prefetch-20260928
Sep 27, 2026
Merged

ValarDragon merged 1 commit into
mainfrom
perf/affine-pending-prefetch-20260928

Conversation

@ValarDragon

@ValarDragon ValarDragon commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Extracted from #507 into an independent PR against main. This contains no
IFMA arithmetic, VK evaluation-plan, IPA coefficient, or decoder changes.

  • Shrink scalar pending affine-addition records from 136 to 104 bytes on
    64-bit Pasta targets. Prescale numerators in the prefix pass to remove a
    multiplication from the dependent backward path without changing the total
    multiplication count.
  • Align prepared-point records to 32 bytes and prefetch two cache-line hints
    16 records ahead, with bounded dense, sparse, and range indexing.
  • Keep the measured exact-field and CPU/OS prefetch policy in a private cached
    no_std detector, so this split does not depend on the IFMA PR. The prefetch
    instruction itself is baseline-safe on x86-64.

No proof, transcript, protocol, or feature changes. Scalar exceptional-case
fallback and failure atomicity are preserved.

Validation

Local ARM, Rust 1.91.0:

  • Batch-affine focused tests: 20 passed.
  • Prefetch/staging and detector tests: 4 passed.
  • Record layout tests: 3 passed.
  • Full prepared zero-check tests: 30 passed, 2 existing timing tests ignored.
  • No-default-feature glv,orbits and multicore configurations checked.
  • Formatting, git diff --check, and independent arithmetic, pointer-bound,
    dependency, and API review passed.

Tests exercise non-unit prefixes, late zero denominators, odd lanes, every
point-preservation guard, dense/sparse/range staging, and detector bit boundaries.
The existing unrelated SqrtTableHelpers dead-code warning remains in the
no-default-feature configuration.

Linux native, Rust 1.97.1, rebased onto the immediate pre-#518 main:

  • Full release Pasta library suite: 252 passed, 3 existing tests ignored.
  • prefetch_detection_matches_standard_library reports prefetch enabled.
  • Serial no-default-feature glv,orbits check passed, with the existing
    unrelated SqrtTableHelpers warning.
  • Every benchmark process independently validated the 64 distinct real proofs
    and rejected altered B1/B2 instances and canonical IPA blinding scalars.

The Pasta prepared record stays 96 bytes, so its table footprint does not grow.

Isolated prover and verifier impact

Compare main f2c0f322 with only this PR applied. The candidate's production
tree is byte-identical to merged #518 (0c771246). No IFMA, IPA/VK, or decoder
changes are included. Both use the same diagnostic-only benchmark.rs
overlay (SHA-256
fa020e7f70f18c7a163e26534427e30b05d6efc57b098fe089af6e149cbbf224).

linux-1, AMD EPYC 9654 VMware guest, pinned CPU 2, one Rayon worker,
Rust 1.97.1 release, default features plus circuit,x86_64-asm,orbits.
Builds used separate fresh target directories and fixed, hashed executables.
The initial shared-target build was rejected before any measurements because
Cargo reused the control binary for the candidate.

Proving uses one A/B/B/A bracket, ten Criterion samples per workload per
process, two-second warmup and 15-second requested measurement. Values below
are pooled medians of times / iters, twenty samples per arm. Key generation,
preparation and fixture construction are outside timing. Prepared proving is
warmed steady-state, not first-proof latency. The prepared and unprepared
harnesses differ in RNG setup; compare each mode against its own control.

Proving workload Main (ms) #518 (ms) Duration reduction
Prepared, one Action 473.865 457.441 3.5%
Prepared, two Actions, padded payment 754.574 730.837 3.1%
Prepared, two real spends 752.650 731.489 2.8%
Prepared, four Actions 1329.309 1288.774 3.0%
Unprepared, two real Actions 888.747 881.269 0.8%

Verification uses two complete A/B/B/A brackets, three warmups and fifteen
samples per size per process, using the same immutable corpus. Each bracket
has thirty samples per arm. Preparation, corpus loading, entry cloning,
positive validation and rejection checks are outside timing. Each proof has
two real Actions, so batch size two means four Actions.

Verifier mode / proofs First bracket reduction Repeat reduction
Prepared / 1 6.3% 5.6%
Prepared / 2 7.4% 7.1%
Prepared / 16 3.4% 4.0%
Prepared / 64 0.6% 0.8%
Unprepared / 1 1.6% 1.6%
Unprepared / 2 1.3% 1.4%
Unprepared / 16 0.9% 1.0%
Unprepared / 64 1.4% 1.5%

All raw samples are retained. The first prepared single-proof control had a
seven-sample scheduling excursion; this motivated the complete verifier
repeat, not selective sample removal. The prepared 64-proof difference is
roughly neutral given VM noise; sub-1% improvements are small observations,
not significance claims. These results do not establish default-feature,
other-hardware, first-proof, or multicore performance.

The two initial control qualifications preceded an interruption by roughly
eight hours. The resumed comparison had fresh idle-host telemetry and fresh
balanced controls; qualification is not claimed to be immediately adjacent.

Frozen diagnostic revisions: control 0eaf9fa4, candidate 0785578b.
Corpus SHA-256:
c8bd4e97159e14cb1bafef13c482c82cc1dcac23efe7cf03d3e8cab3513cfd8e.
Raw data and binary hashes are retained under:

  • /home/valar/benchmarks/pr518-affine-prover-verifier-20260928/logs/linux-1
  • /home/valar/benchmarks/pr518-verifier-repeat-20260928/logs/linux-1

The former combined #507 2x benchmark is not attributed to this PR.

API surface

No function, signature, visibility, trait, enum, or feature additions.
The existing pub(crate) PreparedPoint<F> gains #[repr(align(32))];
Pasta alignment changes from 8 to 32 bytes while the record remains 96 bytes.
Smaller generic-field records can gain padding. The prefetch helpers and
CPU/OS detector remain private.

The corresponding changes were removed from #507. The separate IPA/VK
split is #517; #507 retains its IFMA backend and existing decoder.

@v12-auditor

v12-auditor Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Note

Complete: Audit complete. V12 did not find any issues that need review.

Open the full results here.

Analyzed three files, diff c71526e...f2947c0.

@ValarDragon
ValarDragon force-pushed the perf/affine-pending-prefetch-20260928 branch from f2947c0 to 3ce82f4 Compare September 27, 2026 23:28
@ValarDragon
ValarDragon merged commit 0c77124 into main Sep 27, 2026
70 checks passed
@ValarDragon
ValarDragon deleted the perf/affine-pending-prefetch-20260928 branch September 27, 2026 23:55
lovesh added a commit to PolymeshAssociation/arkworks-algebra that referenced this pull request Sep 28, 2026
Batch-affine Pippenger finishes each level's additions in two passes: the prefix pass scales every numerator by its lane's product of earlier denominators, and the backward pass reads the slope directly and completes the addition. This replaces the separate batch inversion, its per-level prefix-product allocation, the pass that multiplied inverses by numerators, and the serial zero-denominator scan. After Zakura arkworks-rs#518 (zakura-core/common#518). msm_batch_affine on Pallas is 4-6% faster natively from 2^8 bases and 1-3% faster on wasm32.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ValarDragon pushed a commit that referenced this pull request Sep 28, 2026
Bump the workspace version and intra-workspace requirements from 2.0.0
to 2.1.0-rc.0 and assemble the v2.1.0-rc.0 changelog sections, consuming
all 29 pending fragments. This is the first release of zakura-equihash,
imported in #44.

Editorial changes to the assembled entries:
- Consolidate the ten incremental zakura-equihash solver speedup entries
  (#502–#516) into one entry, and give the C-compiler removal from #509
  its own line: the crate ships for the first time, so its section
  describes the delta from upstream equihash 0.3.0 rather than the
  development history.
- Merge the three zakura-bls12-381 final-exponentiation entries (#485,
  #494, #501) into one that states the variable-time inversion and the
  public-input timing contract, and drop the bare "#494" cross-reference.
- Remove technique-only wording from the #483 (bellman), #489
  (bls12_381), and #518 (pasta_curves) entries; #518 claims no speedup,
  so only the reduced scratch memory remains.
- Rephrase the #494 bellman entry in the past tense used elsewhere.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant