Skip to content

About

Scrapy Py docker to run on AWS. This is one step of many to have AWS services consume URLS and crawl urls to populate a shared Open Search index.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

Β 

History

275 Commits

Folders and files

Repository files navigation

Web Crawler for Archived Sites

This project is a Scrapy crawler that runs on a Linux server. It serves as the data collection engine for an aggregated search system.

New to this repo? See QUICKSTART.md to validate your local setup with two of the simplest crawlers before reading further.

πŸ— Architecture

This repo crawls static, archived websites. It pushes each site's converted JSONL to an S3 bucket (see "Push Pipeline" below). One site is not archived: fdrlibrary crawls www.fdrlibrary.org, a live site that its owners still maintain. The client has no access to that site to connect it to the search index directly. An upgrade or redesign of that site will probably break its spider. A downstream Lambda, outside this repo, watches that bucket and indexes into OpenSearch. The search front end queries OpenSearch directly, through Drupal's search_api. Drupal does not trigger or control this repo's crawling or pushing.

What triggers a crawl remains an open question, out of scope for this repo. Options include manual invocation, crontab.example's schedule, or some other interface.

πŸš€ Setup & Installation

Prerequisites

  • The Python version in .python-version
  • AWS credentials, for push and crawl-and-push only (see "Credentials" below)

Local Setup

Run every command in this README from the repository root. The repository root is the directory that contains scrapy.cfg. Do not run commands in the inner archive_crawler/ package directory.

# 1. Create a virtual environment
python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate

# 2. Check the Python version. It must match .python-version.
python --version

# 3. Install dependencies
pip install -r requirements.txt

If python --version does not match .python-version, delete the venv directory. Install the correct Python version, then do these steps again.

requirements.txt pins every package, including indirect dependencies, to a tested version. Do not edit it by hand. See "Dependency Maintenance" below.


πŸ•·οΈ Running Locally (Development)

generic_crawl_harvest and generic_crawl form a two-phase spider pair. Both run entirely locally. They are starter and example tooling, not a production-ready scraper for an arbitrary new site. See the "Step-by-step: generic harvester" section in HARVESTING.md for usage. Each spider's own docstring holds its full -a argument list.


🧭 Nav Harvester Spiders (No Sitemap)

Four spiders have no sitemap to work from. They use NavHarvesterMixin (archive_crawler/spiders/nav_harvest.py) instead β€” one spider per site doing nav link-following, listing-pagination-walking, and content extraction, all in a single crawl:

./scrape_index_pipeline crawl open_obama_whitehouse

Replace with any of: letsmove, obama_whitehouse, trumpwhitehouse.

See ARCHITECTURE.md for the listing-fingerprint mechanism these spiders rely on. obama_whitehouse.py and letsmove.py are its fullest worked example. See the "Step-by-step: nav harvester" section in HARVESTING.md for the full walkthrough. All harvester and content CSVs land under data/{source_site}/. The root data/ directory is git-tracked, through data/.gitkeep. The .csv files themselves are gitignored.


πŸ—Ί Sitemap-Based Archive Spiders

The Clinton (CW1 through CW6), Biden, GWBush whitehouse, and FDR Library spiders discover their own URLs from each site's committed sitemap. Each one scrapes content in the same run, through SitemapUrlSpiderMixin (archive_crawler/spiders/base.py):

./scrape_index_pipeline crawl clintonwhitehouse2

Replace with any of: clintonwhitehouse1, clintonwhitehouse3 through 6, bidenwhitehouse, georgewbush_whitehouse, fdrlibrary. Never pass -O or -o here. See the "Never pass -O/-o to a multi-FEEDS-entry spider" section in ARCHITECTURE.md.

These 9 spiders each model themselves on sitemap_harvest, the generic, one-size-fits-all sitemap URL harvester. It is not part of running any of them. It exists only to explore a new sitemap-based site's URL shape, before writing that site's own spider (see the "Sitemap harvester" section in HARVESTING.md).


⚠️ Warnings Column

no_body, no_title, and short_body are not exclusions. A real page was fetched, so the row stays in the main output CSV. The warnings column flags it instead (comma-separated, if more than one applies):

Warning Meaning
no_body The body selectors returned empty text. full_text/teaser_text are empty strings. title still extracts normally, if present.
short_body Body text was extracted, but it falls under SHORT_BODY_THRESHOLD (default 30 characters). Override it per-spider (class attribute) or per-run (-a short_body_threshold=<N>).
no_title No title could be extracted. title falls back to _slug_title(url): the last URL path segment. The extension gets stripped, and -/_ turn to spaces. No title-casing applies (e.g. pp99-1.html becomes pp99 1). This is a synthesized title, not an authored one. The warnings column is what signals that.

🚫 Exclusion & Dropped Output

Each scrape spider automatically writes up to two CSVs alongside the output CSV, when the spider closes. No -O/-a flag is needed (-a exclusions_file=<path>/-a dropped_file=<path> overrides either derived default). Each row holds the skipped URL and a typed reason. Neither file's rows ever appear in the main output CSV. The split depends on whether a harvest row exists for the URL:

  • {source_site}_exclusions.csv β€” a URL rejected before it was ever a harvest candidate, so it has no harvest row at all. For the 4 no-sitemap spiders, this means a rules:-matched link. Real link-following found it, then dropped it before it was ever requested. For the 9 sitemap-based spiders, this means a sitemap entry that failed the extension allowlist, or matched a rules: entry. Either way, the spider dropped it before writing a harvest row for it. Read this file for a per-rule audit of what got excluded, and why.
  • {source_site}_dropped.csv β€” a URL that already has a harvest row, then got rejected. This happens post-fetch (a bad response) or post-harvest-row (fetched fine, judged non-content).

Two invariants hold for every one of the 13 in-scope sites. scraped + dropped = harvested always holds. For the 9 sitemap-based spiders specifically, harvested + excluded = sitemap total also holds. The 4 no-sitemap spiders have no fixed "total" to reconcile excluded against. A nav crawl's link-discovery has no fixed URL list to bound it, unlike a sitemap.

Reason File Description
url_pattern:/foo/ Exclusions The URL matched a known non-content path prefix.
extension:<ext> Exclusions Sitemap-based spiders only (CW1–6, Biden, GWBush, FDR Library). The sitemap entry failed the site's extension allowlist (e.g. a PDF or image). NavHarvesterMixin-based spiders (all 4 no-sitemap spiders) filter the same way during link-following. They do not log it β€” see "Watch out for" below.
frameset Dropped The page is a frameset with no extractable content.
non_text_response Dropped The response body is not text. Example: a binary file, served from an extension-less URL a link-following crawl swept up.
http_404 Dropped The page returned an HTTP 404.
http_3xx Dropped A redirect went unfollowed (redirects are disabled globally).
http_5xx Dropped The server returned an error.
network_error:<type> Dropped The connection failed at the network level.
search_listing_page Dropped open_obama_whitehouse.py-specific: a /search//search/type/* pagination page, followed only for dataset-link discovery.
pagination_listing_page Dropped PetitionsSpiderMixin-specific: a root or /responses pagination page (?page=N), followed only for petition-link discovery.
listing_page Dropped NavHarvesterMixin-specific (all 4 no-sitemap spiders): the page has a detected listing container (see ARCHITECTURE.md). Its own content goes unscraped.

Watch out for: NavHarvesterMixin-based spiders (the 4 no-sitemap sites) never log a link that _filter_web_urls drops for failing the extension allowlist. That link is silently excluded from following, with no extension:* row anywhere, unlike the 9 sitemap-based spiders' own extension-allowlist rejections during sitemap parsing. _walk_listing_pagination's own pagination-continuation pages (page 2, 3, and on) never get a harvest row either way. A non_text_response logged there does not count toward scraped + dropped = harvested β€” it is diagnostic only. The standalone sitemap_harvest.py exploration tool keeps its own unrelated {source_site}_harvest-dropped.csv (for sitemap-listed URLs that fail its own extension check), separate from everything above.


πŸ“Š URL Gap Analysis

audit_url_gaps.py compares the harvest CSV against the output CSV, and groups unaccounted-for URLs by path prefix:

python audit_url_gaps.py \
  --harvest data/clintonwhitehouse2/clintonwhitehouse2_harvest.csv \
  --output  data/clintonwhitehouse2/clintonwhitehouse2.csv \
  --depth 3 \
  --source-site clintonwhitehouse2

Use --depth 0 to report only the total count, with no path grouping.


βš™οΈ Recommended Run Settings

Run large archives (CW4–6, GWBush) on a remote server. Always launch through scrape_index_pipeline, never a bare scrapy crawl β€” see "Always use the wrapper" below. Override the default throttling with --download-delay/--concurrent-requests-per-domain, not a bare environment variable. settings.py does not read DOWNLOAD_DELAY/CONCURRENT_REQUESTS* from the environment (only FEED_URI, CLOSESPIDER_PAGECOUNT, and DEPTH_LIMIT do). Prefixing the command with DOWNLOAD_DELAY=0.15 ... silently does nothing, and the crawl runs at the settings.py defaults (CONCURRENT_REQUESTS_PER_DOMAIN=4, DOWNLOAD_DELAY=0.25) instead.

The right override on the remote server depends on how many crawls are running there concurrently. The shared constraint is combined outbound load, not any single crawl's own politeness:

concurrent crawls DOWNLOAD_DELAY CONCURRENT_REQUESTS_PER_DOMAIN
1 0.12 10
2 0.15 8
3 0.2 6
4–5 0.25 4 (matches the local default β€” no override needed)
6–7 0.5 2
8+ 1 1
./scrape_index_pipeline crawl georgewbush_whitehouse \
  --download-delay 0.12 \
  --concurrent-requests-per-domain 10

To launch on the remote server itself, SSH in. Background the crawl with nohup/disown, so it survives a disconnect. Point --logfile at a path under that site's data/{site}/ directory, to monitor progress:

ssh user@example-remote-host \
  "cd /home/scrapy/nara-scrapy-crawler && \
   nohup ./scrape_index_pipeline crawl obama_whitehouse \
     --download-delay 0.12 \
     --concurrent-requests-per-domain 10 \
     --logfile data/www.obamawhitehouse/obama_whitehouse-20261231.log \
     > /dev/null 2>&1 & disown"

Launch only one crawl per SSH invocation. Chaining several backgrounded launches together in a single call is unreliable, and can silently drop some of them. The SSH command itself may hang past a client-side timeout, until the entire remote process tree exits, including the disowned job. That is expected, not a stuck connection. Its eventual return is a reliable signal the crawl actually finished.

Before raising throttling further, check the target domain's robots.txt for a Crawl-delay directive. ROBOTSTXT_OBEY = False means Scrapy will not enforce it automatically, so it is easy to run faster than the site operator has asked for, without noticing.

settings.py also sets MEMUSAGE_LIMIT_MB=8192, on the assumption these crawls run on a resource-rich remote server. If a crawl's memory footprint exceeds that limit (for example, a crawler trap on a faceted-search or listing-heavy site generates unbounded unique URLs), Scrapy closes the spider gracefully and flushes the feed export. The OS never gets the chance to OOM-kill the process and lose all buffered output. Override it per-run with --memusage-limit N on crawl/ crawl-and-push (for example, a lower value for local dev testing).


πŸ›‘ Always Use the Wrapper

For every one of the 13 in-scope content spiders, launch through ./scrape_index_pipeline crawl/crawl-and-push, never a bare scrapy crawl <site> call. This is not only a style preference. scrape_index_pipeline's own _crawl step checks the spider process's exit code before continuing. A crawl-and-push run whose crawl exits nonzero never reaches push at all. A bare scrapy crawl, run by hand or scripted outside the wrapper, has no such gate. Nothing stops its output from being pushed later, by a separate push call, with no record of whether that crawl actually finished.

generic_crawl, generic_crawl_harvest, and sitemap_harvest are the exception. All three are one-off exploratory tools for a site not yet onboarded (see HARVESTING.md), outside scrape_index_pipeline's own site registry (archive_crawler/pipeline/registry.py) by design. They have no wrapper equivalent, and running them directly is correct.

πŸ—‚ CSV Naming Convention

All harvester and content output files follow one consistent naming scheme. Every spider writes to its own path automatically, except generic_crawl/generic_crawl_harvest. Those two are one-off exploratory tools with no fixed site identity (see "Running Locally" above), and the only two spiders that still require -O/-o for any output at all. Do not pass -O <path> to any of the 13 in-scope content spiders, to redirect their output. Every one of them has a two-entry custom_settings['FEEDS'] (harvest and content). Scrapy's CLI setting replaces that whole dict, rather than adding to it, which silently drops the harvest CSV and corrupts the content CSV's own shape (see ARCHITECTURE.md). Use -a exclusions_file=<path>/-a dropped_file=<path> for the exclusions/dropped CSVs. For sitemap_harvest specifically, where -O does not apply at all, use -a harvest_file=<path>/-a dropped_file=<path> (its own, unrelated dropped_file).

File Contents
data/{source_site}/{source_site}_harvest.csv One of two automatic FEEDS outputs from the same run: the surviving URL list. This applies to both sitemap-based spiders (CW1–6, Biden, GWBush) and NavHarvesterMixin sites that also extract content (e.g. obama_whitehouse.py, letsmove.py).
data/{source_site}/{source_site}.csv The final content output (includes a warnings column β€” see "Warnings column" above).
data/{source_site}/{source_site}_exclusions.csv Skipped URLs with typed reasons (written when the spider closes).
data/{source_site}/{source_site}-errors-{timestamp}.log The Scrapy ERROR-level log (written by the ErrorFileLogger extension).

Test subsets append -test: {source_site}_harvest-test.csv, {source_site}-test.csv.

{source_site} matches the SOURCE_SITE value in the spider (for example, www.obamawhitehouse, clintonwhitehouse2).


βž• Adding a New Site

See HARVESTING.md for the full process: choosing a harvester type, pre-code discovery, creating either a no-sitemap or sitemap-based spider, and validating the output.


πŸ”Ž Push Pipeline

scrape_index_pipeline takes a site's content CSV through validation, per-site warning-based row filtering, and CSV-to-JSONL conversion, then pushes the result to S3. This project's responsibility ends at that upload. A downstream Lambda watches the bucket, and handles indexing on the OpenSearch side (including any reconciliation against existing index contents). Nothing in this repo deletes or reconciles index contents. Three subcommands:

# Validate/filter/convert/push an existing CSV, no crawl
./scrape_index_pipeline push clintonwhitehouse1

# Run the spider only - scrapy crawl <site>, nothing else
./scrape_index_pipeline crawl clintonwhitehouse1

# Crawl the site first, then do everything push does
./scrape_index_pipeline crawl-and-push clintonwhitehouse1

<site> is either a spider name (bidenwhitehouse) or a source_site (www.bidenwhitehouse) β€” see archive_crawler/pipeline/registry.py. push is the primary path. Per "CSVs are frozen source of truth" (data/8-03/), re-invoking a crawl is the exception, not the default action. Run this from the repo root. The relative data/ paths assume that working directory.

Every command takes exactly one site β€” there is no --all. To run against every site, call the command once per site, the way crontab.example does. Each call gets its own exit code, and one site aborting never blocks the sites after it.

Overrides

Run -h on any subcommand for the full flag list. A few behaviors are worth knowing, that are not obvious from the flag descriptions alone:

  • --csv is push-only. crawl/crawl-and-push never accept a CSV path override. Passing -O to the spider would silently corrupt output. See the "Never pass -O/-o to a multi-FEEDS-entry spider" section in ARCHITECTURE.md. Only the converted JSONL is redirectable after a crawl.
  • --logfile diverts the entire crawl log away from the terminal (Scrapy writes to one or the other, never both). ErrorFileLogger's own ERROR-level file keeps recording regardless. When stdout is a real terminal, a spinner and an elapsed-seconds counter fill the gap this otherwise leaves blank.
  • push/crawl-and-push run four checks before uploading anything, in order: a zero-row CSV, a missing *_dropped.csv, a critical sitemap/pagination row, then an ordinary http_5xx/network_error count against --error-threshold N. See ABORT_CONDITIONS.md for what each one catches and why. --bypass skips only the last of the four. Each site's spider class sets its own default threshold through ERROR_THRESHOLD (see ARCHITECTURE.md), normally 3. Passing --error-threshold on the CLI always overrides that default.

scrape_index_pipeline_interactive prompts for site, mode, and any relevant overrides, instead of requiring them as CLI arguments. It previews every file the run will touch, and confirms before running the equivalent scrape_index_pipeline command β€” the same division of labor as the old run_crawl_interactive.sh and run_crawl.sh pattern this replaces. It is simpler than the bare CLI by design (no --jsonl/--logfile path prompts). Use scrape_index_pipeline directly for finer control.

See ARCHITECTURE.md for what each pipeline module (registry.py/validate.py/filter_rows.py/ convert.py/push.py) actually does.

Credentials

push/crawl-and-push need AWS credentials, and NARA_S3_BUCKET and NARA_ENV set, to upload. NARA_ENV (dev, stage or prod) picks which environment's OpenSearch index the upload reaches. This project uses boto3's own default provider chain as-is. Real AWS_ACCESS_KEY_ID and similar environment variables take priority, when present. Copy .env.example to a gitignored .env, to configure a fallback credentials file or profile, and the target bucket and region, for a server or workstation with no AWS environment variables of its own.


πŸ“‚ Project Structure

Each file's own docstring or comments carry the full detail. This is just a map.

Path What is there
spiders/generic_crawl_harvest.py, spiders/generic_crawl.py The generic two-phase spider pair (see "Running Locally" above)
spiders/base.py ArchiveSpiderMixin, SitemapUrlSpiderMixin, PetitionsSpiderMixin
spiders/nav_harvest.py NavHarvesterMixin β€” see ARCHITECTURE.md
spiders/exclusion_logging.py ExclusionLoggingMixin β€” writes {SOURCE_SITE}_exclusions.csv
spiders/sitemap_harvest.py Generic sitemap onboarding harvester β€” see HARVESTING.md
exclusion_rules.py, exclusion_rules/ Per-domain URL exclusion rules β€” see ARCHITECTURE.md
filter_rules/ Per-source_site push-time warning filter β€” see the "Push pipeline stages" section in ARCHITECTURE.md
items.py ArchiveItem, HarvestItem schemas
extensions/error_log.py ErrorFileLogger
audit_url_gaps.py Post-hoc URL gap analysis tool (see "URL Gap Analysis" above)
pipeline/, scrape_index_pipeline, scrape_index_pipeline_interactive Push pipeline (see "Push Pipeline" above)
requirements.in Direct dependencies. Edit this file, not requirements.txt
requirements.txt Every package pinned to a tested version. pip-compile generates it from requirements.in
.python-version The Python version for every environment
crontab.example Example weekly re-crawl schedule for all 13 sites, 2-parallel-max

πŸ›  Deployment

The crawl server holds a git clone of this repository. To deploy a change:

  1. Merge the change into main.
  2. On the server, in the repository root, run git pull.
  3. If requirements.txt changed, run venv/bin/pip install -r requirements.txt.

The server venv must use the Python version in .python-version.

πŸ“¦ Dependency Maintenance

requirements.in lists the direct dependencies. requirements.txt pins those and all their indirect dependencies to exact versions. Every install gets the same versions, so a new upstream release cannot break a new setup.

A package can remove a private name (a name that starts with _) in any release. An unpinned indirect dependency can thus break the crawler with no change in this repository. For example, w3lib 2.5.0 removed _safe_chars, which Scrapy 2.11.0 imports.

Add or change a dependency

  1. Edit requirements.in.
  2. Run pip install pip-tools, in the venv.
  3. Run pip-compile requirements.in. This rewrites requirements.txt.
  4. Run pip install -r requirements.txt.
  5. Do a short test crawl, then commit both files.

Upgrade dependencies

Run pip-compile --upgrade requirements.in to move every package to its newest allowed version. To upgrade one package only, run pip-compile --upgrade-package <name> requirements.in. Do a test crawl of each spider before you commit an upgrade.

A version limit in requirements.in has a comment that gives its reason. Remove the limit only when that reason no longer applies. For example, w3lib<2.5 can go when Scrapy moves to a version that does not import _safe_chars.

About

Scrapy Py docker to run on AWS. This is one step of many to have AWS services consume URLS and crawl urls to populate a shared Open Search index.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages