Reporting this as an issue per CONTRIBUTING. It was first found by @tallyhuhu in #139, which was labeled bug and ops and then closed only because it had no linked issue. As far as I can tell no issue was opened afterwards, so I am opening one. @tallyhuhu, if you want to pick it back up, it is yours.
What happens
attempt_download resumes with Range: bytes=<part size>- and treats every non 2xx status as a failure (crates/snapshots/src/download.rs:506 on main):
let mut response = request.send().and_then(|r| r.error_for_status())?;
When the .part file already holds the whole archive, that range starts at the end of the resource, and a conforming server answers 416 Range Not Satisfiable with Content-Range: bytes */<size> (RFC 9110, section 15.5.17). As a result:
- all 10 attempts get the same 416, about 45 s goes to backoff, and the restore fails
- the part file and its
.part.url marker stay in place, so every later run repeats this
--force does not help: force_download_and_extract_both calls the same download_archive, and the staging directory is only removed after extraction
- the complete archive on disk is never used; the operator has to find and delete the
.part by hand and download everything again
Reproduction
nginx 1.24 serving a 200000 byte file, using the functions from download.rs on main (2a3e8ab) built against the same reqwest:
| state of the staging dir |
result |
| empty (control) |
downloads, file identical |
90000 byte .part (control) |
resumes, file identical |
complete 200000 byte .part |
10 x 416, fails after 45 s |
| same, second run |
10 more 416s, fails again |
Reachability, stated honestly
The part only ends up complete if the process stops after the last byte is written and before the std::fs::rename at line 556 (Ctrl-C, an OOM kill, a container stop). That window is narrow. The cost when it is hit is not, because nothing short of manual cleanup gets the node out of it.
A part that is longer than the resource gets the same 416 and is stuck the same way, for example if the file behind a URL is republished with a smaller size.
Note on the fix, if you want one
I have a patch ready and tested: on a 416 to a resume request, read the size from Content-Range: bytes */N. If it equals the part size, the part is complete and is promoted as usual. Otherwise the part is removed, so the next attempt starts from byte 0.
Two tests in download.rs cover this, and both fail on main. Each guard is mutation tested: never promoting, always promoting, and keeping the stale part each fail a test. cargo test -p arc-snapshots passes (74 tests), and cargo clippy -p arc-snapshots --all-targets -- -D warnings and cargo fmt --check are clean. Against the same nginx setup, the complete part is promoted after a single 416, and the oversized part is dropped and downloaded again in full, byte identical.
Happy to open a PR if you assign this to me, or to leave the description here if you would rather fix it yourselves.
Reporting this as an issue per CONTRIBUTING. It was first found by @tallyhuhu in #139, which was labeled
bugandopsand then closed only because it had no linked issue. As far as I can tell no issue was opened afterwards, so I am opening one. @tallyhuhu, if you want to pick it back up, it is yours.What happens
attempt_downloadresumes withRange: bytes=<part size>-and treats every non 2xx status as a failure (crates/snapshots/src/download.rs:506onmain):When the
.partfile already holds the whole archive, that range starts at the end of the resource, and a conforming server answers416 Range Not SatisfiablewithContent-Range: bytes */<size>(RFC 9110, section 15.5.17). As a result:.part.urlmarker stay in place, so every later run repeats this--forcedoes not help:force_download_and_extract_bothcalls the samedownload_archive, and the staging directory is only removed after extraction.partby hand and download everything againReproduction
nginx 1.24 serving a 200000 byte file, using the functions from
download.rsonmain(2a3e8ab) built against the same reqwest:.part(control).partReachability, stated honestly
The part only ends up complete if the process stops after the last byte is written and before the
std::fs::renameat line 556 (Ctrl-C, an OOM kill, a container stop). That window is narrow. The cost when it is hit is not, because nothing short of manual cleanup gets the node out of it.A part that is longer than the resource gets the same 416 and is stuck the same way, for example if the file behind a URL is republished with a smaller size.
Note on the fix, if you want one
I have a patch ready and tested: on a 416 to a resume request, read the size from
Content-Range: bytes */N. If it equals the part size, the part is complete and is promoted as usual. Otherwise the part is removed, so the next attempt starts from byte 0.Two tests in
download.rscover this, and both fail onmain. Each guard is mutation tested: never promoting, always promoting, and keeping the stale part each fail a test.cargo test -p arc-snapshotspasses (74 tests), andcargo clippy -p arc-snapshots --all-targets -- -D warningsandcargo fmt --checkare clean. Against the same nginx setup, the complete part is promoted after a single 416, and the oversized part is dropped and downloaded again in full, byte identical.Happy to open a PR if you assign this to me, or to leave the description here if you would rather fix it yourselves.