Skip to content

perf: make quoted-printable decoding beat 0.9.0 on dense escapes too - #264

Merged
kurok merged 1 commit into
masterfrom
perf/qp-dense-escapes
Sep 18, 2026
Merged

kurok merged 1 commit into
masterfrom
perf/qp-dense-escapes

Conversation

@kurok

@kurok kurok commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Comparing the 0.9.0 release against master turned up exactly one benchmark where the release still won: parse_qp_dense_escapes, by 58% on the x86 gate and 20% on an M4. This closes that gap and then some.

Why it was there

#229 replaced the quoted_printable crate with a run-copying decoder and measured −39%. That was true of the benchmarks that existed — but parse_qp_dense_escapes arrived later, with #223, so the one input the new decoder is worst at was never compared against the code it replaced.

Measured directly on =C3=A9 × 20000 (120 KB, an escape every three bytes):

dense ordinary QP
the crate 0.9.0 used 166 µs 188 µs
ours, before 236 µs 118 µs

The two changes

Both in vendor/mailparse/src/qp.rs.

1. Look under the cursor before reaching for memchr. After an escape the next byte is very often another = — every non-ASCII character encodes as two or three consecutive escapes — so memchr was called only to report a match at offset 0, paying its SIMD setup each time. On the dense fixture that was 40,000 calls. This is why the fix helps ordinary mail too, not just the pathological case.

2. Find the first dropped byte a chunk at a time. The rule-1 pre-scan used position, which cannot vectorise because it must stop at the first hit. It now reduces 32-byte chunks with |= and only falls back to a byte-wise scan inside the chunk that failed. is_kept is respelled as arithmetic (b - 0x20 < 0x5F, b - 9 < 2, b == 0x0D) so that reduction is branchless. On 120 KB with nothing to drop — nearly every real body — 63 µs → 5.5 µs.

Results

Local interleaved A/B, Apple M4, 3 rounds, controls within 3.1%:

benchmark master this PR
parse_qp_dense_escapes 0.171 ms 0.102 ms −40%
parse_qp_message 0.129 ms 0.085 ms −34%

Everything else is inside the noise floor. Against 0.9.0 the dense case is now 1.40× faster instead of 1.20× slower, and ordinary quoted-printable 2.6× faster.

Output is unchanged

is_kept was respelled, so the risk is that the set changed. is_kept_is_the_same_set checks all 256 bytes against the original matches! spelling. On top of that the existing crate-agreement tests still pass — the generated corpus at every length, all 65,536 =xy byte pairs, the per-rule cases and the real fixture — and qp_agreement fuzzed 7.3M executions against the crate with no disagreement.

upstream.patch is regenerated (check_vendored_mailparse.sh passes) and PATCH.md records both changes and why they are shaped that way, so the next mailparse hand-merge does not quietly drop them.

Comparing the 0.9.0 release against master turned up exactly one benchmark
where the release still won: parse_qp_dense_escapes, by 58% on the x86 gate
and 20% on an M4.

#229 replaced the quoted_printable crate with a run-copying decoder and
measured -39%. That was true of the benchmarks that existed. The dense shape
did not have one: parse_qp_dense_escapes arrived later with #223, so the one
input the new decoder is worst at was never compared against the code it
replaced. Measured directly, ours took 236 us against the crate's 166 us.

Two changes, both in the vendored qp.rs:

decode_line now looks at the byte under the cursor before calling memchr.
After an escape the next byte is very often another `=` -- every non-ASCII
character encodes as two or three consecutive escapes -- so memchr was being
called only to report a match at offset 0, paying its SIMD setup each time.
On the dense fixture that was 40,000 calls.

The rule-1 pre-scan finds the first dropped byte a chunk at a time rather
than with `position`, which cannot vectorise because it must stop at the
first hit. `is_kept` is respelled as arithmetic so the chunk reduction is
branchless; is_kept_is_the_same_set checks all 256 bytes against the original
spelling. On 120 KB with nothing to drop -- nearly every real body -- that
scan goes from 63 us to 5.5 us.

Local interleaved A/B, M4, 3 rounds, controls within 3.1%:
parse_qp_dense_escapes 0.171 -> 0.102 ms (-40%), parse_qp_message
0.129 -> 0.085 ms (-34%), everything else inside the floor. Against 0.9.0 the
dense case is now 1.40x faster instead of 1.20x slower, and ordinary
quoted-printable 2.6x faster.

Output is unchanged. The vendored crate-agreement tests pass, including all
65,536 `=xy` byte pairs, and qp_agreement fuzzed 7.3M executions against the
crate with no disagreement. upstream.patch is regenerated and PATCH.md
records both changes and why they are shaped that way.

Signed-off-by: kurok <22548029+kurok@users.noreply.github.com>
@kurok
kurok merged commit ed1c55b into master Sep 18, 2026
15 checks passed
@kurok
kurok deleted the perf/qp-dense-escapes branch September 18, 2026 13:55
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant