web-scraper-rust
A polite, single-binary Rust web scraper that crawls pages, converts HTML to Markdown, and tries its best to turn Office documents into something readable too.
It respects robots.txt. It waits between requests. It hashes long URLs into filenames because filesystems have feelings too. It does not do cookies, JavaScript rendering, or anything fancy. If you need a headless browser, this is not your tool.
What It Does
- Crawls internal links from a starting URL (BFS queue)
- Converts HTML pages to Markdown via
htmd - Detects and converts documents (PDF, DOCX, XLSX, PPTX, ODT, RTF, EPUB, CSV...) to Markdown via
anydoc - Respects
robots.txtwith wildcard pattern matching - Deduplicates URLs, normalizes fragments and trailing query params
- Configurable request delay, retry with exponential backoff
- File type filtering via
--include-types/--exclude-types - Path scoping (stay within a subdirectory) or full-site crawl
- Subdomain crawling opt-in
- Single-page mode (grab one page, don't crawl further)
- Structured error logging to JSONL, kept in a separate
_logsdirectory
Tech Stack
| Thing | Choice | Why |
|---|---|---|
| Language | Rust 2024 | Single static binary. No runtime. No GC. No regrets. |
| HTTP client | reqwest + tokio | Async, fast, well-maintained. |
| HTML parsing | scraper | CSS selectors on server-side HTML. |
| HTML→MD | htmd | Strips nav, footer, script tags. Produces clean Markdown. |
| Document conversion | anydoc | Detects format from bytes, converts to Markdown. Falls back to raw. |
| CLI | clap | Derive macros, typed args, good help text. |
| Logging | env_logger | Simple, sufficient. |
Usage
# Scrape a single page, save Markdown
web-scraper https://example.com/page --single
# Crawl entire site, only PDFs and HTML
web-scraper https://example.com --types pdf,html
# Crawl a subdirectory only, exclude images
web-scraper https://example.com/docs --exclude-types png,jpg,gif
# Include subdomains, custom delay
web-scraper https://example.com --subdomains --delay-ms 500
Output goes to ./<domain>_scraped/, logs to ./<domain>_scraped_logs/.
CLI Flags
START_URL Starting URL to scrape
-o, --output DIR Output directory (default: <domain>_scraped)
--subdomains Also crawl subdomains of the base host
--single Only scrape the start URL, don't crawl
--no-scope Crawl any path on the same host
--no-doc-conversion Save documents as-is, skip conversion
--delay-ms N Delay between requests in ms (default: 1000)
--types a,b,c Only save these file types (HTML always crawled)
--exclude-types a,b Skip these file types
--types and --exclude-types are mutually exclusive. Both still crawl HTML pages for links regardless.
Build
just build # clippy + fmt + release build
# or
cargo build --release
Release profile is optimized for size: opt-level = "z", LTO, single codegen unit, stripped.
Error Logs
Errors are logged to scrape_errors.jsonl in the logs directory. Each entry includes timestamp, URL, error type, message, and HTTP status code if applicable. The log file is cleared at the start of each run.
Built to scrape things. Respects robots.txt. Does not require a DevOps team.