Files
2026-08-27 10:46:08 +01:00

84 lines
3.3 KiB
Markdown

# web-scraper-rust
A polite, single-binary Rust web scraper that crawls pages, converts HTML to Markdown, and tries its best to turn Office documents into something readable too.
It respects `robots.txt`. It waits between requests. It hashes long URLs into filenames because filesystems have feelings too. It does not do cookies, JavaScript rendering, or anything fancy. If you need a headless browser, this is not your tool.
## What It Does
- Crawls internal links from a starting URL (BFS queue)
- Converts HTML pages to Markdown via `htmd`
- Detects and converts documents (PDF, DOCX, XLSX, PPTX, ODT, RTF, EPUB, CSV...) to Markdown via `anydoc`
- Respects `robots.txt` with wildcard pattern matching
- Deduplicates URLs, normalizes fragments and trailing query params
- Configurable request delay, retry with exponential backoff
- File type filtering via `--include-types` / `--exclude-types`
- Path scoping (stay within a subdirectory) or full-site crawl
- Subdomain crawling opt-in
- Single-page mode (grab one page, don't crawl further)
- Structured error logging to JSONL, kept in a separate `_logs` directory
## Tech Stack
| Thing | Choice | Why |
|---|---|---|
| Language | Rust 2024 | Single static binary. No runtime. No GC. No regrets. |
| HTTP client | reqwest + tokio | Async, fast, well-maintained. |
| HTML parsing | scraper | CSS selectors on server-side HTML. |
| HTML→MD | htmd | Strips nav, footer, script tags. Produces clean Markdown. |
| Document conversion | anydoc | Detects format from bytes, converts to Markdown. Falls back to raw. |
| CLI | clap | Derive macros, typed args, good help text. |
| Logging | env_logger | Simple, sufficient. |
## Usage
```bash
# Scrape a single page, save Markdown
web-scraper https://example.com/page --single
# Crawl entire site, only PDFs and HTML
web-scraper https://example.com --types pdf,html
# Crawl a subdirectory only, exclude images
web-scraper https://example.com/docs --exclude-types png,jpg,gif
# Include subdomains, custom delay
web-scraper https://example.com --subdomains --delay-ms 500
```
Output goes to `./<domain>_scraped/`, logs to `./<domain>_scraped_logs/`.
## CLI Flags
```
START_URL Starting URL to scrape
-o, --output DIR Output directory (default: <domain>_scraped)
--subdomains Also crawl subdomains of the base host
--single Only scrape the start URL, don't crawl
--no-scope Crawl any path on the same host
--no-doc-conversion Save documents as-is, skip conversion
--delay-ms N Delay between requests in ms (default: 1000)
--types a,b,c Only save these file types (HTML always crawled)
--exclude-types a,b Skip these file types
```
`--types` and `--exclude-types` are mutually exclusive. Both still crawl HTML pages for links regardless.
## Build
```bash
just build # clippy + fmt + release build
# or
cargo build --release
```
Release profile is optimized for size: `opt-level = "z"`, LTO, single codegen unit, stripped.
## Error Logs
Errors are logged to `scrape_errors.jsonl` in the logs directory. Each entry includes timestamp, URL, error type, message, and HTTP status code if applicable. The log file is cleared at the start of each run.
---
**Built to scrape things. Respects robots.txt. Does not require a DevOps team.**