Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 10 additions & 1 deletion .github/workflows/test.yml
Original file line number Diff line number Diff line change
Expand Up @@ -264,9 +264,18 @@ jobs:

- name: Fuzz
run: |
# `-a` turns on debug assertions, which is what makes the deferred
# modes' retention checkable at all (#239): a leaf keeps offsets into
# the buffer it was parsed from, and `Retained::of` falls back to a copy
# when the bytes it is handed are not in that buffer. The fallback is
# deliberate -- it cannot read out of bounds -- but it is silent, so
# without the assertion a slip in the offsets degrades to the copying
# this issue removed and nothing says so. The deep-fuzz workflow stays
# on the release build, where more executions per second matter more
# than this one invariant.
for target in $FUZZ_TARGETS; do
echo "::group::$target"
cargo fuzz run "$target" -- \
cargo fuzz run -a "$target" -- \
-max_total_time="$FUZZ_SECONDS" \
-seed=1 \
-print_final_stats=1
Expand Down
24 changes: 24 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,6 +9,30 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

### Changed

- **Lazy modes borrow the payload instead of copying every deferred part** (#239). Deferring
a decode used to copy the part's encoded bytes out of the message, which is the opposite
of what deferring is for: parsing a 96 MiB single-attachment message in `mode="lazy"` cost
96 MiB *on top of* the payload the caller still held, and base64 is 1.33x what it encodes,
so a retained part could cost more than the decoded bytes it avoided producing. A deferred
part now keeps its offsets in the buffer it was parsed from, and the result keeps that
buffer alive. Measured on a 96 MiB message with one unread attachment: peak RSS 192.3 MiB
before, 96.2 MiB after. Counted in the core, on the attachment-heavy fixture: a lazy tree
peaked at 798,111 bytes and now peaks at 14,332 -- the same as a metadata tree, because a
range is not a copy -- and a flat lazy parse dropped from 790,993 to 23,398.

**This changes the memory contract, and is the reason to read this entry.** A `PyLazyMail`
attachment or a `PyLazyMimePart` leaf pins the payload it was parsed from for as long as it
is reachable, so keeping one attachment out of a mailbox keeps that whole message rather
than just the part. The pin is per message even in `parse_many`, so one slot never holds
another's payload. A caller who wants the bytes without the message reads `content`, which
is a decoded copy, and drops the attachment. Nothing about the values changes: every mode
returns what it returned, including for a message whose header block had to be repaired
(#150), where the offsets index the rebuilt copy that now travels with the result.

One exception, invisible from Python: the leaves *inside* a `message/rfc822` node still hold
copies. Their bytes were produced by decoding that node's body, so they are in no caller
buffer to point into.

- **The parsing core is a crate, and has Rust tests for the first time** (#236). It was a
`#[path]`-included file: the binding declared it as a module, and both fuzz targets
reached across the tree to include the same source again under different cfg. Nothing
Expand Down
23 changes: 17 additions & 6 deletions Readme.md
Original file line number Diff line number Diff line change
Expand Up @@ -487,10 +487,14 @@ runner; the ratios are what to read.

Two things to know before choosing it:

**It trades memory for decoding.** The encoded bytes of every attachment are
retained until the message is dropped, and base64 is about 1.33x the size of what
it encodes — so a retained part costs *more* than the decoded bytes it avoids
producing. Right for one attachment out of twenty; wrong for all twenty.
**It pins the payload.** A deferred attachment does not copy itself out of the
message; it remembers where it sits in the payload you passed in, and keeps that
payload alive for as long as it is reachable. So the result holds one message's
worth of memory rather than two — parsing a 96 MiB message now costs 96 MiB and
used to cost 192 MiB — but the payload is not released when you drop your own
reference to it. Keeping one attachment out of a mailbox keeps that message, not
its attachment. If you want the bytes without the message, read `content`, which
is a decoded copy, and drop the attachment.

**A `DecodeError` moves.** A part whose `Content-Transfer-Encoding` cannot be
decoded fails the whole parse in full mode, and fails on `content` here — so a
Expand Down Expand Up @@ -574,8 +578,9 @@ data = pdf.content # decoded here, and only this one

`mode="metadata"` decodes nothing *and retains nothing*: a node reports
`encoded_size` in place of `content` and there is no way to ask for the bytes.
That is the difference between the two — lazy mode keeps a copy of every leaf so
it can decode one later, metadata mode keeps none and is the cheaper sweep.
That is the difference between the two — a lazy leaf remembers where it sits in
the payload so it can decode itself later, which keeps that payload alive;
metadata mode remembers nothing and is the sweep that holds nothing.

On the attachment-heavy fixture (767 KiB), median of three interleaved rounds on
the CI runner: full tree 0.556 ms, `mode="lazy"` with nothing read 0.077 ms,
Expand All @@ -595,6 +600,12 @@ has to deliver eagerly. Its node therefore arrives with `is_decoded` already
`True`, and unlike `parse_email(mode="metadata")` a deferred tree *can* raise
`DecodeError` for such a part. Nothing else is decoded.

It is also the one place a lazy leaf holds bytes of its own. The leaves *inside*
an embedded message were parsed out of that decoded body rather than out of your
payload, so they keep copies; everywhere else a leaf keeps offsets. The
difference is not visible from Python — a leaf decodes the same way either way —
but it is why a lazy tree of a bounce is not quite free.

`walk` accepts a node from any mode and yields nodes of the same type.

#### Which API when
Expand Down
2 changes: 2 additions & 0 deletions bench/src/main.rs
Original file line number Diff line number Diff line change
Expand Up @@ -92,12 +92,14 @@ fn main() {
"tree-lazy" => Box::new(|| {
mail_parser::parse_tree_deferred(&payload, true)
.expect("the fixture must parse")
.root
.children
.len()
}),
"tree-metadata" => Box::new(|| {
mail_parser::parse_tree_deferred(&payload, false)
.expect("the fixture must parse")
.root
.children
.len()
}),
Expand Down
39 changes: 33 additions & 6 deletions bench/tests/allocs.rs
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,13 @@
//! sweep does not decode bodies, lazy mode so an untouched attachment is not
//! copied. Those claims have been prose. This counts them.
//!
//! `peak` is the number to read for lazy mode. Since #239 a deferred leaf keeps
//! offsets into the buffer it was parsed from rather than a copy of itself, so
//! what a lazy result holds no longer scales with the bodies it deferred -- and
//! the table is where that stops being a claim. What it cannot show is the
//! payload the caller still holds, which those offsets now pin; the Python side
//! measures that, because the pin lives in the binding.
//!
//! Exact equality against a committed table, not bounds. A bound absorbs a
//! regression silently until it crosses the bound; an exact number turns any
//! change in allocation behaviour into a diff someone has to look at and either
Expand Down Expand Up @@ -190,17 +197,37 @@ fn the_modes_allocate_what_they_promise() {
"an untouched lazy parse must hold less than a full one \
(lazy {lazy:?}, full {full:?})"
);
// Deliberately asserted only here. Lazy mode retains each part's *encoded*
// bytes, and base64 is 4/3 of what it decodes to -- so on a small
// attachment-bearing message a lazy tree genuinely holds MORE than a decoded
// one (measured: 6696 against 6323 bytes on attachment_message.eml). The
// trade only pays once the bodies are large, which is the case the mode was
// written for and the only case where this ordering is a promise.
assert!(
tree_lazy.peak < tree.peak,
"on a message with real attachments, a lazy tree must hold less than a \
decoded one (lazy {tree_lazy:?}, decoded {tree:?})"
);
// The #239 claim, and the one worth guarding: what a deferred leaf retains is
// a range, so a lazy tree's footprint is its structure and not its bodies.
// Until #239 this assertion was false by a wide margin -- a lazy tree of this
// fixture peaked at 798,111 bytes against the 14,332 it peaks at now -- and
// on a small attachment-bearing message a lazy tree held *more* than a fully
// decoded one, because base64 is 4/3 of what it decodes to and the mode kept
// the encoded side.
assert!(
tree_lazy.peak < payload.len() / 4,
"a lazy tree peaked at {} bytes on a {} byte message; it retains offsets \
rather than bytes, so it should be far under a quarter of the input",
tree_lazy.peak,
payload.len()
);
// The flat mode's version of the same claim. Weaker on purpose: flat lazy
// decodes every body part, so its floor is the bodies, and only the
// attachments are deferred. A quarter of a full parse is the margin that
// holds for a message whose weight is in its attachments -- which is the
// message this mode is for.
assert!(
lazy.peak * 4 < full.peak,
"an untouched lazy parse held {} bytes against a full parse's {}; the \
attachments it deferred should not be in either number",
lazy.peak,
full.peak
);
assert!(
metadata.peak < payload.len() / 4,
"metadata mode peaked at {} bytes on a {} byte message; it decodes no \
Expand Down
Loading