# web-scraper-rust A polite, single-binary Rust web scraper that crawls pages, converts HTML to Markdown, and tries its best to turn Office documents into something readable too. It respects `robots.txt`. It waits between requests. It hashes long URLs into filenames because filesystems have feelings too. It does not do cookies, JavaScript rendering, or anything fancy. If you need a headless browser, this is not your tool. ## What It Does - Crawls internal links from a starting URL (BFS queue) - Converts HTML pages to Markdown via `htmd` - Detects and converts documents (PDF, DOCX, XLSX, PPTX, ODT, RTF, EPUB, CSV...) to Markdown via `anydoc` - Respects `robots.txt` with wildcard pattern matching - Deduplicates URLs, normalizes fragments and trailing query params - Configurable request delay, retry with exponential backoff - File type filtering via `--include-types` / `--exclude-types` - Path scoping (stay within a subdirectory) or full-site crawl - Subdomain crawling opt-in - Single-page mode (grab one page, don't crawl further) - Structured error logging to JSONL, kept in a separate `_logs` directory ## Tech Stack | Thing | Choice | Why | |---|---|---| | Language | Rust 2024 | Single static binary. No runtime. No GC. No regrets. | | HTTP client | reqwest + tokio | Async, fast, well-maintained. | | HTML parsing | scraper | CSS selectors on server-side HTML. | | HTML→MD | htmd | Strips nav, footer, script tags. Produces clean Markdown. | | Document conversion | anydoc | Detects format from bytes, converts to Markdown. Falls back to raw. | | CLI | clap | Derive macros, typed args, good help text. | | Logging | env_logger | Simple, sufficient. | ## Usage ```bash # Scrape a single page, save Markdown web-scraper https://example.com/page --single # Crawl entire site, only PDFs and HTML web-scraper https://example.com --types pdf,html # Crawl a subdirectory only, exclude images web-scraper https://example.com/docs --exclude-types png,jpg,gif # Include subdomains, custom delay web-scraper https://example.com --subdomains --delay-ms 500 ``` Output goes to `./_scraped/`, logs to `./_scraped_logs/`. ## CLI Flags ``` START_URL Starting URL to scrape -o, --output DIR Output directory (default: _scraped) --subdomains Also crawl subdomains of the base host --single Only scrape the start URL, don't crawl --no-scope Crawl any path on the same host --no-doc-conversion Save documents as-is, skip conversion --delay-ms N Delay between requests in ms (default: 1000) --types a,b,c Only save these file types (HTML always crawled) --exclude-types a,b Skip these file types ``` `--types` and `--exclude-types` are mutually exclusive. Both still crawl HTML pages for links regardless. ## Build ```bash just build # clippy + fmt + release build # or cargo build --release ``` Release profile is optimized for size: `opt-level = "z"`, LTO, single codegen unit, stripped. ## Error Logs Errors are logged to `scrape_errors.jsonl` in the logs directory. Each entry includes timestamp, URL, error type, message, and HTTP status code if applicable. The log file is cleared at the start of each run. --- **Built to scrape things. Respects robots.txt. Does not require a DevOps team.**