Files
2026-08-27 10:46:08 +01:00

3.3 KiB

web-scraper-rust

A polite, single-binary Rust web scraper that crawls pages, converts HTML to Markdown, and tries its best to turn Office documents into something readable too.

It respects robots.txt. It waits between requests. It hashes long URLs into filenames because filesystems have feelings too. It does not do cookies, JavaScript rendering, or anything fancy. If you need a headless browser, this is not your tool.

What It Does

  • Crawls internal links from a starting URL (BFS queue)
  • Converts HTML pages to Markdown via htmd
  • Detects and converts documents (PDF, DOCX, XLSX, PPTX, ODT, RTF, EPUB, CSV...) to Markdown via anydoc
  • Respects robots.txt with wildcard pattern matching
  • Deduplicates URLs, normalizes fragments and trailing query params
  • Configurable request delay, retry with exponential backoff
  • File type filtering via --include-types / --exclude-types
  • Path scoping (stay within a subdirectory) or full-site crawl
  • Subdomain crawling opt-in
  • Single-page mode (grab one page, don't crawl further)
  • Structured error logging to JSONL, kept in a separate _logs directory

Tech Stack

Thing Choice Why
Language Rust 2024 Single static binary. No runtime. No GC. No regrets.
HTTP client reqwest + tokio Async, fast, well-maintained.
HTML parsing scraper CSS selectors on server-side HTML.
HTML→MD htmd Strips nav, footer, script tags. Produces clean Markdown.
Document conversion anydoc Detects format from bytes, converts to Markdown. Falls back to raw.
CLI clap Derive macros, typed args, good help text.
Logging env_logger Simple, sufficient.

Usage

# Scrape a single page, save Markdown
web-scraper https://example.com/page --single

# Crawl entire site, only PDFs and HTML
web-scraper https://example.com --types pdf,html

# Crawl a subdirectory only, exclude images
web-scraper https://example.com/docs --exclude-types png,jpg,gif

# Include subdomains, custom delay
web-scraper https://example.com --subdomains --delay-ms 500

Output goes to ./<domain>_scraped/, logs to ./<domain>_scraped_logs/.

CLI Flags

START_URL              Starting URL to scrape
-o, --output DIR       Output directory (default: <domain>_scraped)
    --subdomains       Also crawl subdomains of the base host
    --single           Only scrape the start URL, don't crawl
    --no-scope         Crawl any path on the same host
    --no-doc-conversion  Save documents as-is, skip conversion
    --delay-ms N       Delay between requests in ms (default: 1000)
    --types a,b,c      Only save these file types (HTML always crawled)
    --exclude-types a,b  Skip these file types

--types and --exclude-types are mutually exclusive. Both still crawl HTML pages for links regardless.

Build

just build      # clippy + fmt + release build
# or
cargo build --release

Release profile is optimized for size: opt-level = "z", LTO, single codegen unit, stripped.

Error Logs

Errors are logged to scrape_errors.jsonl in the logs directory. Each entry includes timestamp, URL, error type, message, and HTTP status code if applicable. The log file is cleared at the start of each run.


Built to scrape things. Respects robots.txt. Does not require a DevOps team.