This project is a Scrapy crawler that runs on a Linux server. It serves as the data collection engine for an aggregated search system.
New to this repo? See QUICKSTART.md to validate your local setup with two of the simplest crawlers before reading further.
This repo crawls static, archived websites. It pushes each site's converted JSONL to an S3 bucket (see "Push Pipeline" below). One site is not archived: fdrlibrary crawls www.fdrlibrary.org, a live site that its owners still maintain. The client has no access to that site to connect it to the search index directly. An upgrade or redesign of that site will probably break its spider. A downstream Lambda, outside this repo, watches that bucket and indexes into OpenSearch. The search front end queries OpenSearch directly, through Drupal's search_api. Drupal does not trigger or control this repo's crawling or pushing.
What triggers a crawl remains an open question, out of scope for this repo. Options include manual invocation, crontab.example's schedule, or some other interface.
- The Python version in
.python-version - AWS credentials, for
pushandcrawl-and-pushonly (see "Credentials" below)
Run every command in this README from the repository root. The repository root is the directory that contains scrapy.cfg. Do not run commands in the inner archive_crawler/ package directory.
# 1. Create a virtual environment
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# 2. Check the Python version. It must match .python-version.
python --version
# 3. Install dependencies
pip install -r requirements.txtIf python --version does not match .python-version, delete the venv directory. Install the correct Python version, then do these steps again.
requirements.txt pins every package, including indirect dependencies, to a tested version. Do not edit it by hand. See "Dependency Maintenance" below.
generic_crawl_harvest and generic_crawl form a two-phase spider pair. Both run entirely locally. They are starter and example tooling, not a production-ready scraper for an arbitrary new site. See the "Step-by-step: generic harvester" section in HARVESTING.md for usage. Each spider's own docstring holds its full -a argument list.
Four spiders have no sitemap to work from. They use
NavHarvesterMixin (archive_crawler/spiders/nav_harvest.py) instead β one
spider per site doing nav link-following, listing-pagination-walking, and
content extraction, all in a single crawl:
./scrape_index_pipeline crawl open_obama_whitehouseReplace with any of: letsmove, obama_whitehouse, trumpwhitehouse.
See ARCHITECTURE.md
for the listing-fingerprint mechanism these spiders rely on. obama_whitehouse.py
and letsmove.py are its fullest worked example. See the "Step-by-step:
nav harvester" section in HARVESTING.md for the full walkthrough. All harvester and
content CSVs land under data/{source_site}/. The root data/ directory is
git-tracked, through data/.gitkeep. The .csv files themselves are gitignored.
The Clinton (CW1 through CW6), Biden, GWBush whitehouse, and FDR Library
spiders discover their own URLs from each site's committed sitemap. Each one scrapes content
in the same run, through SitemapUrlSpiderMixin (archive_crawler/spiders/base.py):
./scrape_index_pipeline crawl clintonwhitehouse2Replace with any of: clintonwhitehouse1, clintonwhitehouse3 through 6,
bidenwhitehouse, georgewbush_whitehouse, fdrlibrary. Never pass -O or -o here.
See the "Never pass -O/-o to a multi-FEEDS-entry
spider" section in ARCHITECTURE.md.
These 9 spiders each model themselves on sitemap_harvest, the generic,
one-size-fits-all sitemap URL harvester. It is not part of running any
of them. It exists only to explore a new sitemap-based site's URL
shape, before writing that site's own spider (see the "Sitemap
harvester" section in HARVESTING.md).
no_body, no_title, and short_body are not exclusions. A real
page was fetched, so the row stays in the main output CSV. The
warnings column flags it instead (comma-separated, if more than one
applies):
| Warning | Meaning |
|---|---|
no_body |
The body selectors returned empty text. full_text/teaser_text are empty strings. title still extracts normally, if present. |
short_body |
Body text was extracted, but it falls under SHORT_BODY_THRESHOLD (default 30 characters). Override it per-spider (class attribute) or per-run (-a short_body_threshold=<N>). |
no_title |
No title could be extracted. title falls back to _slug_title(url): the last URL path segment. The extension gets stripped, and -/_ turn to spaces. No title-casing applies (e.g. pp99-1.html becomes pp99 1). This is a synthesized title, not an authored one. The warnings column is what signals that. |
Each scrape spider automatically writes up to two CSVs alongside the
output CSV, when the spider closes. No -O/-a flag is needed
(-a exclusions_file=<path>/-a dropped_file=<path> overrides either
derived default). Each row holds the skipped URL and a typed reason.
Neither file's rows ever appear in the main output CSV. The split
depends on whether a harvest row exists for the URL:
{source_site}_exclusions.csvβ a URL rejected before it was ever a harvest candidate, so it has no harvest row at all. For the 4 no-sitemap spiders, this means arules:-matched link. Real link-following found it, then dropped it before it was ever requested. For the 9 sitemap-based spiders, this means a sitemap entry that failed the extension allowlist, or matched arules:entry. Either way, the spider dropped it before writing a harvest row for it. Read this file for a per-rule audit of what got excluded, and why.{source_site}_dropped.csvβ a URL that already has a harvest row, then got rejected. This happens post-fetch (a bad response) or post-harvest-row (fetched fine, judged non-content).
Two invariants hold for every one of the 13 in-scope sites. scraped + dropped = harvested always holds. For the 9 sitemap-based spiders specifically, harvested + excluded = sitemap total also holds. The 4 no-sitemap spiders have no fixed "total" to reconcile excluded against. A nav crawl's link-discovery has no fixed URL list to bound it, unlike a sitemap.
| Reason | File | Description |
|---|---|---|
url_pattern:/foo/ |
Exclusions | The URL matched a known non-content path prefix. |
extension:<ext> |
Exclusions | Sitemap-based spiders only (CW1β6, Biden, GWBush, FDR Library). The sitemap entry failed the site's extension allowlist (e.g. a PDF or image). NavHarvesterMixin-based spiders (all 4 no-sitemap spiders) filter the same way during link-following. They do not log it β see "Watch out for" below. |
frameset |
Dropped | The page is a frameset with no extractable content. |
non_text_response |
Dropped | The response body is not text. Example: a binary file, served from an extension-less URL a link-following crawl swept up. |
http_404 |
Dropped | The page returned an HTTP 404. |
http_3xx |
Dropped | A redirect went unfollowed (redirects are disabled globally). |
http_5xx |
Dropped | The server returned an error. |
network_error:<type> |
Dropped | The connection failed at the network level. |
search_listing_page |
Dropped | open_obama_whitehouse.py-specific: a /search//search/type/* pagination page, followed only for dataset-link discovery. |
pagination_listing_page |
Dropped | PetitionsSpiderMixin-specific: a root or /responses pagination page (?page=N), followed only for petition-link discovery. |
listing_page |
Dropped | NavHarvesterMixin-specific (all 4 no-sitemap spiders): the page has a detected listing container (see ARCHITECTURE.md). Its own content goes unscraped. |
Watch out for: NavHarvesterMixin-based spiders (the 4 no-sitemap
sites) never log a link that _filter_web_urls drops for failing the
extension allowlist. That link is silently excluded from following,
with no extension:* row anywhere, unlike the 9 sitemap-based spiders'
own extension-allowlist rejections during sitemap parsing.
_walk_listing_pagination's own pagination-continuation pages (page 2,
3, and on) never get a harvest row either way. A non_text_response
logged there does not count toward scraped + dropped = harvested β it
is diagnostic only. The standalone sitemap_harvest.py exploration
tool keeps its own unrelated {source_site}_harvest-dropped.csv (for
sitemap-listed URLs that fail its own extension check), separate from
everything above.
audit_url_gaps.py compares the harvest CSV against the output CSV, and groups unaccounted-for URLs by path prefix:
python audit_url_gaps.py \
--harvest data/clintonwhitehouse2/clintonwhitehouse2_harvest.csv \
--output data/clintonwhitehouse2/clintonwhitehouse2.csv \
--depth 3 \
--source-site clintonwhitehouse2Use --depth 0 to report only the total count, with no path grouping.
Run large archives (CW4β6, GWBush) on a remote server. Always launch
through scrape_index_pipeline, never a bare scrapy crawl β see
"Always use the wrapper" below. Override the default throttling with
--download-delay/--concurrent-requests-per-domain, not a bare
environment variable. settings.py does not read
DOWNLOAD_DELAY/CONCURRENT_REQUESTS* from the environment (only
FEED_URI, CLOSESPIDER_PAGECOUNT, and DEPTH_LIMIT do). Prefixing
the command with DOWNLOAD_DELAY=0.15 ... silently does nothing, and
the crawl runs at the settings.py defaults
(CONCURRENT_REQUESTS_PER_DOMAIN=4, DOWNLOAD_DELAY=0.25) instead.
The right override on the remote server depends on how many crawls are running there concurrently. The shared constraint is combined outbound load, not any single crawl's own politeness:
| concurrent crawls | DOWNLOAD_DELAY |
CONCURRENT_REQUESTS_PER_DOMAIN |
|---|---|---|
| 1 | 0.12 | 10 |
| 2 | 0.15 | 8 |
| 3 | 0.2 | 6 |
| 4β5 | 0.25 | 4 (matches the local default β no override needed) |
| 6β7 | 0.5 | 2 |
| 8+ | 1 | 1 |
./scrape_index_pipeline crawl georgewbush_whitehouse \
--download-delay 0.12 \
--concurrent-requests-per-domain 10To launch on the remote server itself, SSH in. Background the crawl
with nohup/disown, so it survives a disconnect. Point --logfile
at a path under that site's data/{site}/ directory, to monitor
progress:
ssh user@example-remote-host \
"cd /home/scrapy/nara-scrapy-crawler && \
nohup ./scrape_index_pipeline crawl obama_whitehouse \
--download-delay 0.12 \
--concurrent-requests-per-domain 10 \
--logfile data/www.obamawhitehouse/obama_whitehouse-20261231.log \
> /dev/null 2>&1 & disown"Launch only one crawl per SSH invocation. Chaining several backgrounded launches together in a single call is unreliable, and can silently drop some of them. The SSH command itself may hang past a client-side timeout, until the entire remote process tree exits, including the disowned job. That is expected, not a stuck connection. Its eventual return is a reliable signal the crawl actually finished.
Before raising throttling further, check the target domain's
robots.txt for a Crawl-delay directive. ROBOTSTXT_OBEY = False
means Scrapy will not enforce it automatically, so it is easy to run
faster than the site operator has asked for, without noticing.
settings.py also sets MEMUSAGE_LIMIT_MB=8192, on the assumption
these crawls run on a resource-rich remote server. If a crawl's memory
footprint exceeds that limit (for example, a crawler trap on a
faceted-search or listing-heavy site generates unbounded unique URLs),
Scrapy closes the spider gracefully and flushes the feed export. The OS
never gets the chance to OOM-kill the process and lose all buffered
output. Override it per-run with --memusage-limit N on crawl/
crawl-and-push (for example, a lower value for local dev testing).
For every one of the 13 in-scope content spiders, launch through
./scrape_index_pipeline crawl/crawl-and-push, never a bare scrapy crawl <site> call. This is not only a style preference.
scrape_index_pipeline's own _crawl step checks the spider process's
exit code before continuing. A crawl-and-push run whose crawl exits
nonzero never reaches push at all. A bare scrapy crawl, run by
hand or scripted outside the wrapper, has no such gate. Nothing stops
its output from being pushed later, by a separate push call, with no
record of whether that crawl actually finished.
generic_crawl, generic_crawl_harvest, and sitemap_harvest are the
exception. All three are one-off exploratory tools for a site not yet
onboarded (see HARVESTING.md), outside scrape_index_pipeline's own
site registry (archive_crawler/pipeline/registry.py) by design. They
have no wrapper equivalent, and running them directly is correct.
All harvester and content output files follow one consistent naming
scheme. Every spider writes to its own path automatically, except
generic_crawl/generic_crawl_harvest. Those two are one-off
exploratory tools with no fixed site identity (see "Running Locally"
above), and the only two spiders that still require -O/-o for any
output at all. Do not pass -O <path> to any of the 13 in-scope
content spiders, to redirect their output. Every one of them has a
two-entry custom_settings['FEEDS'] (harvest and content). Scrapy's
CLI setting replaces that whole dict, rather than adding to it, which
silently drops the harvest CSV and corrupts the content CSV's own shape
(see ARCHITECTURE.md). Use -a exclusions_file=<path>/-a dropped_file=<path>
for the exclusions/dropped CSVs. For sitemap_harvest specifically,
where -O does not apply at all, use -a harvest_file=<path>/-a dropped_file=<path>
(its own, unrelated dropped_file).
| File | Contents |
|---|---|
data/{source_site}/{source_site}_harvest.csv |
One of two automatic FEEDS outputs from the same run: the surviving URL list. This applies to both sitemap-based spiders (CW1β6, Biden, GWBush) and NavHarvesterMixin sites that also extract content (e.g. obama_whitehouse.py, letsmove.py). |
data/{source_site}/{source_site}.csv |
The final content output (includes a warnings column β see "Warnings column" above). |
data/{source_site}/{source_site}_exclusions.csv |
Skipped URLs with typed reasons (written when the spider closes). |
data/{source_site}/{source_site}-errors-{timestamp}.log |
The Scrapy ERROR-level log (written by the ErrorFileLogger extension). |
Test subsets append -test: {source_site}_harvest-test.csv, {source_site}-test.csv.
{source_site} matches the SOURCE_SITE value in the spider (for example, www.obamawhitehouse, clintonwhitehouse2).
See HARVESTING.md for the full process: choosing a harvester type, pre-code discovery, creating either a no-sitemap or sitemap-based spider, and validating the output.
scrape_index_pipeline takes a site's content CSV through validation,
per-site warning-based row filtering, and CSV-to-JSONL conversion, then
pushes the result to S3. This project's responsibility ends at that
upload. A downstream Lambda watches the bucket, and handles indexing
on the OpenSearch side (including any reconciliation against existing
index contents). Nothing in this repo deletes or reconciles index
contents. Three subcommands:
# Validate/filter/convert/push an existing CSV, no crawl
./scrape_index_pipeline push clintonwhitehouse1
# Run the spider only - scrapy crawl <site>, nothing else
./scrape_index_pipeline crawl clintonwhitehouse1
# Crawl the site first, then do everything push does
./scrape_index_pipeline crawl-and-push clintonwhitehouse1<site> is either a spider name (bidenwhitehouse) or a source_site
(www.bidenwhitehouse) β see archive_crawler/pipeline/registry.py.
push is the primary path. Per "CSVs are frozen source of truth"
(data/8-03/), re-invoking a crawl is the exception, not the default
action. Run this from the repo root. The relative data/ paths assume
that working directory.
Every command takes exactly one site β there is no --all. To run
against every site, call the command once per site, the way
crontab.example does. Each call gets its own exit code, and one site
aborting never blocks the sites after it.
Run -h on any subcommand for the full flag list. A few behaviors are
worth knowing, that are not obvious from the flag descriptions alone:
--csvispush-only.crawl/crawl-and-pushnever accept a CSV path override. Passing-Oto the spider would silently corrupt output. See the "Never pass-O/-oto a multi-FEEDS-entry spider" section in ARCHITECTURE.md. Only the converted JSONL is redirectable after a crawl.--logfilediverts the entire crawl log away from the terminal (Scrapy writes to one or the other, never both).ErrorFileLogger's own ERROR-level file keeps recording regardless. When stdout is a real terminal, a spinner and an elapsed-seconds counter fill the gap this otherwise leaves blank.push/crawl-and-pushrun four checks before uploading anything, in order: a zero-row CSV, a missing*_dropped.csv, a critical sitemap/pagination row, then an ordinaryhttp_5xx/network_errorcount against--error-threshold N. See ABORT_CONDITIONS.md for what each one catches and why.--bypassskips only the last of the four. Each site's spider class sets its own default threshold throughERROR_THRESHOLD(see ARCHITECTURE.md), normally 3. Passing--error-thresholdon the CLI always overrides that default.
scrape_index_pipeline_interactive prompts for site, mode, and any
relevant overrides, instead of requiring them as CLI arguments. It
previews every file the run will touch, and confirms before running the
equivalent scrape_index_pipeline command β the same division of labor
as the old run_crawl_interactive.sh and run_crawl.sh pattern this
replaces. It is simpler than the bare CLI by design (no --jsonl/--logfile
path prompts). Use scrape_index_pipeline directly for finer control.
See ARCHITECTURE.md
for what each pipeline module (registry.py/validate.py/filter_rows.py/
convert.py/push.py) actually does.
push/crawl-and-push need AWS credentials, and NARA_S3_BUCKET and
NARA_ENV set, to upload. NARA_ENV (dev, stage or prod) picks
which environment's OpenSearch index the upload reaches. This project
uses boto3's own default provider chain as-is.
Real AWS_ACCESS_KEY_ID and similar environment variables take
priority, when present. Copy .env.example to a
gitignored .env, to configure a fallback credentials file or profile,
and the target bucket and region, for a server or workstation with no
AWS environment variables of its own.
Each file's own docstring or comments carry the full detail. This is just a map.
| Path | What is there |
|---|---|
spiders/generic_crawl_harvest.py, spiders/generic_crawl.py |
The generic two-phase spider pair (see "Running Locally" above) |
spiders/base.py |
ArchiveSpiderMixin, SitemapUrlSpiderMixin, PetitionsSpiderMixin |
spiders/nav_harvest.py |
NavHarvesterMixin β see ARCHITECTURE.md |
spiders/exclusion_logging.py |
ExclusionLoggingMixin β writes {SOURCE_SITE}_exclusions.csv |
spiders/sitemap_harvest.py |
Generic sitemap onboarding harvester β see HARVESTING.md |
exclusion_rules.py, exclusion_rules/ |
Per-domain URL exclusion rules β see ARCHITECTURE.md |
filter_rules/ |
Per-source_site push-time warning filter β see the "Push pipeline stages" section in ARCHITECTURE.md |
items.py |
ArchiveItem, HarvestItem schemas |
extensions/error_log.py |
ErrorFileLogger |
audit_url_gaps.py |
Post-hoc URL gap analysis tool (see "URL Gap Analysis" above) |
pipeline/, scrape_index_pipeline, scrape_index_pipeline_interactive |
Push pipeline (see "Push Pipeline" above) |
requirements.in |
Direct dependencies. Edit this file, not requirements.txt |
requirements.txt |
Every package pinned to a tested version. pip-compile generates it from requirements.in |
.python-version |
The Python version for every environment |
crontab.example |
Example weekly re-crawl schedule for all 13 sites, 2-parallel-max |
The crawl server holds a git clone of this repository. To deploy a change:
- Merge the change into
main. - On the server, in the repository root, run
git pull. - If
requirements.txtchanged, runvenv/bin/pip install -r requirements.txt.
The server venv must use the Python version in .python-version.
requirements.in lists the direct dependencies. requirements.txt pins those and all their indirect dependencies to exact versions. Every install gets the same versions, so a new upstream release cannot break a new setup.
A package can remove a private name (a name that starts with _) in any release. An unpinned indirect dependency can thus break the crawler with no change in this repository. For example, w3lib 2.5.0 removed _safe_chars, which Scrapy 2.11.0 imports.
- Edit
requirements.in. - Run
pip install pip-tools, in the venv. - Run
pip-compile requirements.in. This rewritesrequirements.txt. - Run
pip install -r requirements.txt. - Do a short test crawl, then commit both files.
Run pip-compile --upgrade requirements.in to move every package to its newest allowed version. To upgrade one package only, run pip-compile --upgrade-package <name> requirements.in. Do a test crawl of each spider before you commit an upgrade.
A version limit in requirements.in has a comment that gives its reason. Remove the limit only when that reason no longer applies. For example, w3lib<2.5 can go when Scrapy moves to a version that does not import _safe_chars.