From 33ef5251d58d3057e9fe0c944de53930dbc2e50d Mon Sep 17 00:00:00 2001 From: Bartal Laearsson Date: Thu, 27 Aug 2026 10:46:08 +0100 Subject: [PATCH] readem --- README.md | 83 +++++++++++++++++++++++++++++++++++++++++++++++++++++++ 1 file changed, 83 insertions(+) create mode 100644 README.md diff --git a/README.md b/README.md new file mode 100644 index 0000000..5bf2cad --- /dev/null +++ b/README.md @@ -0,0 +1,83 @@ +# web-scraper-rust + +A polite, single-binary Rust web scraper that crawls pages, converts HTML to Markdown, and tries its best to turn Office documents into something readable too. + +It respects `robots.txt`. It waits between requests. It hashes long URLs into filenames because filesystems have feelings too. It does not do cookies, JavaScript rendering, or anything fancy. If you need a headless browser, this is not your tool. + +## What It Does + +- Crawls internal links from a starting URL (BFS queue) +- Converts HTML pages to Markdown via `htmd` +- Detects and converts documents (PDF, DOCX, XLSX, PPTX, ODT, RTF, EPUB, CSV...) to Markdown via `anydoc` +- Respects `robots.txt` with wildcard pattern matching +- Deduplicates URLs, normalizes fragments and trailing query params +- Configurable request delay, retry with exponential backoff +- File type filtering via `--include-types` / `--exclude-types` +- Path scoping (stay within a subdirectory) or full-site crawl +- Subdomain crawling opt-in +- Single-page mode (grab one page, don't crawl further) +- Structured error logging to JSONL, kept in a separate `_logs` directory + +## Tech Stack + +| Thing | Choice | Why | +|---|---|---| +| Language | Rust 2024 | Single static binary. No runtime. No GC. No regrets. | +| HTTP client | reqwest + tokio | Async, fast, well-maintained. | +| HTML parsing | scraper | CSS selectors on server-side HTML. | +| HTML→MD | htmd | Strips nav, footer, script tags. Produces clean Markdown. | +| Document conversion | anydoc | Detects format from bytes, converts to Markdown. Falls back to raw. | +| CLI | clap | Derive macros, typed args, good help text. | +| Logging | env_logger | Simple, sufficient. | + +## Usage + +```bash +# Scrape a single page, save Markdown +web-scraper https://example.com/page --single + +# Crawl entire site, only PDFs and HTML +web-scraper https://example.com --types pdf,html + +# Crawl a subdirectory only, exclude images +web-scraper https://example.com/docs --exclude-types png,jpg,gif + +# Include subdomains, custom delay +web-scraper https://example.com --subdomains --delay-ms 500 +``` + +Output goes to `./_scraped/`, logs to `./_scraped_logs/`. + +## CLI Flags + +``` +START_URL Starting URL to scrape +-o, --output DIR Output directory (default: _scraped) + --subdomains Also crawl subdomains of the base host + --single Only scrape the start URL, don't crawl + --no-scope Crawl any path on the same host + --no-doc-conversion Save documents as-is, skip conversion + --delay-ms N Delay between requests in ms (default: 1000) + --types a,b,c Only save these file types (HTML always crawled) + --exclude-types a,b Skip these file types +``` + +`--types` and `--exclude-types` are mutually exclusive. Both still crawl HTML pages for links regardless. + +## Build + +```bash +just build # clippy + fmt + release build +# or +cargo build --release +``` + +Release profile is optimized for size: `opt-level = "z"`, LTO, single codegen unit, stripped. + +## Error Logs + +Errors are logged to `scrape_errors.jsonl` in the logs directory. Each entry includes timestamp, URL, error type, message, and HTTP status code if applicable. The log file is cleared at the start of each run. + +--- + +**Built to scrape things. Respects robots.txt. Does not require a DevOps team.**