Skip to content

arc-snapshots: a fully downloaded .part file can never be resumed, and --force does not recover #444

Description

@erkancamli

Reporting this as an issue per CONTRIBUTING. It was first found by @tallyhuhu in #139, which was labeled bug and ops and then closed only because it had no linked issue. As far as I can tell no issue was opened afterwards, so I am opening one. @tallyhuhu, if you want to pick it back up, it is yours.

What happens

attempt_download resumes with Range: bytes=<part size>- and treats every non 2xx status as a failure (crates/snapshots/src/download.rs:506 on main):

let mut response = request.send().and_then(|r| r.error_for_status())?;

When the .part file already holds the whole archive, that range starts at the end of the resource, and a conforming server answers 416 Range Not Satisfiable with Content-Range: bytes */<size> (RFC 9110, section 15.5.17). As a result:

  • all 10 attempts get the same 416, about 45 s goes to backoff, and the restore fails
  • the part file and its .part.url marker stay in place, so every later run repeats this
  • --force does not help: force_download_and_extract_both calls the same download_archive, and the staging directory is only removed after extraction
  • the complete archive on disk is never used; the operator has to find and delete the .part by hand and download everything again

Reproduction

nginx 1.24 serving a 200000 byte file, using the functions from download.rs on main (2a3e8ab) built against the same reqwest:

state of the staging dir result
empty (control) downloads, file identical
90000 byte .part (control) resumes, file identical
complete 200000 byte .part 10 x 416, fails after 45 s
same, second run 10 more 416s, fails again

Reachability, stated honestly

The part only ends up complete if the process stops after the last byte is written and before the std::fs::rename at line 556 (Ctrl-C, an OOM kill, a container stop). That window is narrow. The cost when it is hit is not, because nothing short of manual cleanup gets the node out of it.

A part that is longer than the resource gets the same 416 and is stuck the same way, for example if the file behind a URL is republished with a smaller size.

Note on the fix, if you want one

I have a patch ready and tested: on a 416 to a resume request, read the size from Content-Range: bytes */N. If it equals the part size, the part is complete and is promoted as usual. Otherwise the part is removed, so the next attempt starts from byte 0.

Two tests in download.rs cover this, and both fail on main. Each guard is mutation tested: never promoting, always promoting, and keeping the stale part each fail a test. cargo test -p arc-snapshots passes (74 tests), and cargo clippy -p arc-snapshots --all-targets -- -D warnings and cargo fmt --check are clean. Against the same nginx setup, the complete part is promoted after a single 416, and the oversized part is dropped and downloaded again in full, byte identical.

Happy to open a PR if you assign this to me, or to leave the description here if you would rather fix it yourselves.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    tracked internallyThis issue is already tracked internally by the Arc team.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions