Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 29 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,35 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

### Fixed

- **Quoted-printable decoding was slower than 0.9.0 on a body that is mostly escapes.**
Comparing the 0.9.0 release against master turned up one benchmark where the release won:
`parse_qp_dense_escapes`, by 58% on the x86 gate. #229 replaced the `quoted_printable`
crate with a run-copying decoder and measured a 39% win, but the dense shape had no
benchmark then -- `parse_qp_dense_escapes` arrived later, with #223 -- so the one input
the new decoder is worst at was never compared against the code it replaced. Two changes,
both in `vendor/mailparse/src/qp.rs`:

- `decode_line` looks at the byte under the cursor before calling `memchr`. After an
escape the next byte is very often another `=` -- every non-ASCII character encodes as
two or three consecutive escapes -- so `memchr` was being called to be told the match
was at offset 0, paying its SIMD setup each time. On the dense fixture that happened
40,000 times.
- The rule-1 pre-scan finds the first dropped byte a chunk at a time instead of with
`position`, which cannot vectorise because it has to stop at the first hit. `is_kept`
is respelled as arithmetic so the chunk reduction has no branches in it;
`is_kept_is_the_same_set` checks all 256 bytes against the original spelling. On 120 KB
with nothing to drop -- nearly every real body -- that scan goes from 63 us to 5.5 us.

Measured on an Apple M4, 3 interleaved rounds, controls within 3.1%:
`parse_qp_dense_escapes` **0.171 -> 0.102 ms (-40%)** and `parse_qp_message`
**0.129 -> 0.085 ms (-34%)**, with every other benchmark inside the noise floor. Against
0.9.0 the dense case is now 1.40x faster rather than 1.20x slower, and the ordinary
quoted-printable case 2.6x faster. Output is unchanged: the vendored suite's
crate-agreement tests still pass, including every `=xy` byte pair, and `qp_agreement`
ran 7.3M executions against the crate with no disagreement.

### Changed

- **Internal: the binding layer is split into modules, with shared getters and batch-slot
Expand Down
19 changes: 19 additions & 0 deletions vendor/mailparse/PATCH.md
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,25 @@ what the code it replaces returned.
encoded-word path in `src/header.rs` still uses the crate -- header-sized inputs, and
different semantics (`_` -> space, trailing-whitespace restore).

Two things in it are shaped by measurement rather than taste, and both were found by
comparing against the crate this replaced on a body that is *mostly* escapes -- the
shape the original rewrite never had a benchmark for, and on which it decoded 42%
slower than the code it replaced:

- `decode_line` looks at the byte under the cursor before reaching for `memchr`. After
an escape the next byte is very often another `=`, because every non-ASCII character
encodes as two or three consecutive escapes, and calling `memchr` to be told the
match is at offset 0 pays its SIMD setup for nothing. On the dense shape that call
happened 40,000 times. With the check, decoding is faster than the crate on that
shape *and* ~20% faster on ordinary mail.
- `is_kept` is written as arithmetic (`b - 0x20 < 0x5F`, `b - 9 < 2`, `b == 0x0D`)
rather than a `matches!` pattern, so that `first_dropped`'s chunk reduction has no
branches and LLVM can vectorise it. `position` cannot be vectorised -- it has to stop
at the first hit -- so the rule-1 pre-scan walks 32-byte chunks reduced with `|=` and
only falls back to a byte-wise scan inside the chunk that failed. On 120 KB with no
dropped bytes, which is nearly every real body, that scan went from 63 us to 5.5 us.
`is_kept_is_the_same_set` checks all 256 bytes against the original spelling.

`diff -r` against the registry copy shows exactly `src/bytescan.rs`, `src/qp.rs`,
`src/lib.rs` (two `mod` lines and one function), `src/body.rs` (two functions and a
`mod tests`), `Cargo.toml` and this file.
Expand Down
74 changes: 66 additions & 8 deletions vendor/mailparse/src/qp.rs
Original file line number Diff line number Diff line change
Expand Up @@ -37,9 +37,45 @@
use memchr::memchr;

/// True for the bytes the crate's `filter_map` keeps.
///
/// Written as arithmetic rather than `matches!` so that `first_dropped`'s chunk
/// reduction has no branches in it: `b - 0x20 < 0x5F` is `0x20..=0x7E` and
/// `b - 9 < 2` is TAB and LF. The set is identical -- `is_kept_is_the_same_set`
/// checks all 256 bytes against the original spelling -- but the branchy version
/// stopped LLVM vectorising the scan, and that scan is 40% of decoding a body
/// with no dropped bytes in it at all.
#[inline]
fn is_kept(byte: u8) -> bool {
matches!(byte, b'\t' | b'\r' | b'\n' | b' '..=b'~')
byte.wrapping_sub(0x20) < 0x5F || byte.wrapping_sub(9) < 2 || byte == b'\r'
}

/// Offset of the first byte rule 1 drops, or `None` when there is none.
///
/// A chunk at a time, because `position` has to stop at the first hit and so
/// cannot be vectorised, while a fixed-size chunk reduced with `|=` has no early
/// exit and can be. The byte-wise scan then runs only inside the one chunk that
/// failed. Nearly every real body keeps every byte, which is the case this makes
/// fast: 120 KB scans in 5.5 us against 63 us for the `position` loop.
fn first_dropped(input: &[u8]) -> Option<usize> {
const CHUNK: usize = 32;
let mut base = 0;
for chunk in input.chunks_exact(CHUNK) {
let mut dropped = 0u8;
for &byte in chunk {
dropped |= !is_kept(byte) as u8;
}
if dropped != 0 {
return chunk
.iter()
.position(|&byte| !is_kept(byte))
.map(|i| base + i);
}
base += CHUNK;
}
input[base..]
.iter()
.position(|&byte| !is_kept(byte))
.map(|i| base + i)
}

/// Trailing bytes a line is trimmed of. After the filter these are the only
Expand All @@ -53,7 +89,7 @@ fn is_trimmed(byte: u8) -> bool {
/// `input` without the bytes rule 1 drops. Returns `None` when there are none,
/// which is the common case and saves the copy.
fn filter_dropped(input: &[u8]) -> Option<Vec<u8>> {
let first = input.iter().position(|&b| !is_kept(b))?;
let first = first_dropped(input)?;

let mut out = Vec::with_capacity(input.len());
out.extend_from_slice(&input[..first]);
Expand Down Expand Up @@ -88,14 +124,25 @@ fn hex_value(byte: u8) -> Option<u8> {
fn decode_line(line: &[u8], out: &mut Vec<u8>) -> bool {
let mut pos = 0;
while pos < line.len() {
let Some(offset) = memchr(b'=', &line[pos..]) else {
out.extend_from_slice(&line[pos..]);
return true;
// Look under the cursor first. After an escape the next byte is very
// often another `=` -- every non-ASCII character encodes as two or three
// consecutive escapes -- and calling `memchr` to be told the match is at
// offset 0 pays its SIMD setup for nothing. A body that is mostly escapes
// made that call 40,000 times and decoded 42% slower than the crate this
// replaced; with the check it is faster than the crate on that shape and
// 20% faster on ordinary mail too.
let eq = if line[pos] == b'=' {
pos
} else {
let Some(offset) = memchr(b'=', &line[pos..]) else {
out.extend_from_slice(&line[pos..]);
return true;
};
let eq = pos + offset;
out.extend_from_slice(&line[pos..eq]);
eq
};

let eq = pos + offset;
out.extend_from_slice(&line[pos..eq]);

match line.get(eq + 1) {
// `=` is the last byte: a soft break. Nothing is emitted and the next
// line joins this one directly.
Expand Down Expand Up @@ -224,6 +271,16 @@ mod tests {
out
}

/// `is_kept` was respelled as arithmetic to let `first_dropped` vectorise.
/// This is the proof that it is the same set, byte for byte.
#[test]
fn is_kept_is_the_same_set() {
for byte in 0u8..=255 {
let original = matches!(byte, b'\t' | b'\r' | b'\n' | b' '..=b'~');
assert_eq!(super::is_kept(byte), original, "byte {byte:#04x}");
}
}

#[test]
fn decode_robust_matches_the_crate() {
for input in corpus() {
Expand Down Expand Up @@ -302,3 +359,4 @@ mod tests {
assert_agrees(raw);
}
}

86 changes: 72 additions & 14 deletions vendor/mailparse/upstream.patch
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@
# (then re-add this header)
diff -ruN -x target mailparse-0.16.1/Cargo.toml vendored/Cargo.toml
--- mailparse-0.16.1/Cargo.toml 1970-01-01 01:00:01.000000000 +0100
+++ vendored/Cargo.toml 2026-09-17 12:03:55.686053778 +0100
+++ vendored/Cargo.toml 2026-09-18 14:41:56.600854881 +0100
@@ -37,6 +37,12 @@
license = "0BSD"
repository = "https://github.com/staktrace/mailparse"
Expand Down Expand Up @@ -53,7 +53,7 @@ diff -ruN -x target mailparse-0.16.1/Cargo.toml vendored/Cargo.toml
version = "0.17.0"
diff -ruN -x target mailparse-0.16.1/src/body.rs vendored/src/body.rs
--- mailparse-0.16.1/src/body.rs 2006-07-24 02:21:28.000000000 +0100
+++ vendored/src/body.rs 2026-09-17 12:03:55.687864996 +0100
+++ vendored/src/body.rs 2026-09-18 14:41:56.602471471 +0100
@@ -136,20 +136,66 @@
}
}
Expand Down Expand Up @@ -289,7 +289,7 @@ diff -ruN -x target mailparse-0.16.1/src/body.rs vendored/src/body.rs
+}
diff -ruN -x target mailparse-0.16.1/src/bytescan.rs vendored/src/bytescan.rs
--- mailparse-0.16.1/src/bytescan.rs 1970-01-01 01:00:00.000000000 +0100
+++ vendored/src/bytescan.rs 2026-09-17 12:03:55.688675668 +0100
+++ vendored/src/bytescan.rs 2026-09-18 14:41:56.602908597 +0100
@@ -0,0 +1,153 @@
+//! Byte scans that look at a machine word at a time instead of a byte at a time.
+//!
Expand Down Expand Up @@ -446,7 +446,7 @@ diff -ruN -x target mailparse-0.16.1/src/bytescan.rs vendored/src/bytescan.rs
+}
diff -ruN -x target mailparse-0.16.1/src/lib.rs vendored/src/lib.rs
--- mailparse-0.16.1/src/lib.rs 2006-07-24 02:21:28.000000000 +0100
+++ vendored/src/lib.rs 2026-09-17 12:03:55.688290249 +0100
+++ vendored/src/lib.rs 2026-09-18 14:41:56.602639888 +0100
@@ -13,10 +13,12 @@

mod addrparse;
Expand Down Expand Up @@ -495,8 +495,8 @@ diff -ruN -x target mailparse-0.16.1/src/lib.rs vendored/src/lib.rs
#[test]
diff -ruN -x target mailparse-0.16.1/src/qp.rs vendored/src/qp.rs
--- mailparse-0.16.1/src/qp.rs 1970-01-01 01:00:00.000000000 +0100
+++ vendored/src/qp.rs 2026-09-17 12:03:55.687672912 +0100
@@ -0,0 +1,304 @@
+++ vendored/src/qp.rs 2026-09-18 14:41:56.602333178 +0100
@@ -0,0 +1,362 @@
+//! Quoted-printable body decoding, a run at a time instead of a byte at a time.
+//!
+//! This replaces `quoted_printable::decode(body, ParseMode::Robust)` for message
Expand Down Expand Up @@ -536,9 +536,45 @@ diff -ruN -x target mailparse-0.16.1/src/qp.rs vendored/src/qp.rs
+use memchr::memchr;
+
+/// True for the bytes the crate's `filter_map` keeps.
+///
+/// Written as arithmetic rather than `matches!` so that `first_dropped`'s chunk
+/// reduction has no branches in it: `b - 0x20 < 0x5F` is `0x20..=0x7E` and
+/// `b - 9 < 2` is TAB and LF. The set is identical -- `is_kept_is_the_same_set`
+/// checks all 256 bytes against the original spelling -- but the branchy version
+/// stopped LLVM vectorising the scan, and that scan is 40% of decoding a body
+/// with no dropped bytes in it at all.
+#[inline]
+fn is_kept(byte: u8) -> bool {
+ matches!(byte, b'\t' | b'\r' | b'\n' | b' '..=b'~')
+ byte.wrapping_sub(0x20) < 0x5F || byte.wrapping_sub(9) < 2 || byte == b'\r'
+}
+
+/// Offset of the first byte rule 1 drops, or `None` when there is none.
+///
+/// A chunk at a time, because `position` has to stop at the first hit and so
+/// cannot be vectorised, while a fixed-size chunk reduced with `|=` has no early
+/// exit and can be. The byte-wise scan then runs only inside the one chunk that
+/// failed. Nearly every real body keeps every byte, which is the case this makes
+/// fast: 120 KB scans in 5.5 us against 63 us for the `position` loop.
+fn first_dropped(input: &[u8]) -> Option<usize> {
+ const CHUNK: usize = 32;
+ let mut base = 0;
+ for chunk in input.chunks_exact(CHUNK) {
+ let mut dropped = 0u8;
+ for &byte in chunk {
+ dropped |= !is_kept(byte) as u8;
+ }
+ if dropped != 0 {
+ return chunk
+ .iter()
+ .position(|&byte| !is_kept(byte))
+ .map(|i| base + i);
+ }
+ base += CHUNK;
+ }
+ input[base..]
+ .iter()
+ .position(|&byte| !is_kept(byte))
+ .map(|i| base + i)
+}
+
+/// Trailing bytes a line is trimmed of. After the filter these are the only
Expand All @@ -552,7 +588,7 @@ diff -ruN -x target mailparse-0.16.1/src/qp.rs vendored/src/qp.rs
+/// `input` without the bytes rule 1 drops. Returns `None` when there are none,
+/// which is the common case and saves the copy.
+fn filter_dropped(input: &[u8]) -> Option<Vec<u8>> {
+ let first = input.iter().position(|&b| !is_kept(b))?;
+ let first = first_dropped(input)?;
+
+ let mut out = Vec::with_capacity(input.len());
+ out.extend_from_slice(&input[..first]);
Expand Down Expand Up @@ -587,14 +623,25 @@ diff -ruN -x target mailparse-0.16.1/src/qp.rs vendored/src/qp.rs
+fn decode_line(line: &[u8], out: &mut Vec<u8>) -> bool {
+ let mut pos = 0;
+ while pos < line.len() {
+ let Some(offset) = memchr(b'=', &line[pos..]) else {
+ out.extend_from_slice(&line[pos..]);
+ return true;
+ // Look under the cursor first. After an escape the next byte is very
+ // often another `=` -- every non-ASCII character encodes as two or three
+ // consecutive escapes -- and calling `memchr` to be told the match is at
+ // offset 0 pays its SIMD setup for nothing. A body that is mostly escapes
+ // made that call 40,000 times and decoded 42% slower than the crate this
+ // replaced; with the check it is faster than the crate on that shape and
+ // 20% faster on ordinary mail too.
+ let eq = if line[pos] == b'=' {
+ pos
+ } else {
+ let Some(offset) = memchr(b'=', &line[pos..]) else {
+ out.extend_from_slice(&line[pos..]);
+ return true;
+ };
+ let eq = pos + offset;
+ out.extend_from_slice(&line[pos..eq]);
+ eq
+ };
+
+ let eq = pos + offset;
+ out.extend_from_slice(&line[pos..eq]);
+
+ match line.get(eq + 1) {
+ // `=` is the last byte: a soft break. Nothing is emitted and the next
+ // line joins this one directly.
Expand Down Expand Up @@ -723,6 +770,16 @@ diff -ruN -x target mailparse-0.16.1/src/qp.rs vendored/src/qp.rs
+ out
+ }
+
+ /// `is_kept` was respelled as arithmetic to let `first_dropped` vectorise.
+ /// This is the proof that it is the same set, byte for byte.
+ #[test]
+ fn is_kept_is_the_same_set() {
+ for byte in 0u8..=255 {
+ let original = matches!(byte, b'\t' | b'\r' | b'\n' | b' '..=b'~');
+ assert_eq!(super::is_kept(byte), original, "byte {byte:#04x}");
+ }
+ }
+
+ #[test]
+ fn decode_robust_matches_the_crate() {
+ for input in corpus() {
Expand Down Expand Up @@ -801,3 +858,4 @@ diff -ruN -x target mailparse-0.16.1/src/qp.rs vendored/src/qp.rs
+ assert_agrees(raw);
+ }
+}
+
Loading