readem
This commit is contained in:
@@ -0,0 +1,83 @@
|
||||
# web-scraper-rust
|
||||
|
||||
A polite, single-binary Rust web scraper that crawls pages, converts HTML to Markdown, and tries its best to turn Office documents into something readable too.
|
||||
|
||||
It respects `robots.txt`. It waits between requests. It hashes long URLs into filenames because filesystems have feelings too. It does not do cookies, JavaScript rendering, or anything fancy. If you need a headless browser, this is not your tool.
|
||||
|
||||
## What It Does
|
||||
|
||||
- Crawls internal links from a starting URL (BFS queue)
|
||||
- Converts HTML pages to Markdown via `htmd`
|
||||
- Detects and converts documents (PDF, DOCX, XLSX, PPTX, ODT, RTF, EPUB, CSV...) to Markdown via `anydoc`
|
||||
- Respects `robots.txt` with wildcard pattern matching
|
||||
- Deduplicates URLs, normalizes fragments and trailing query params
|
||||
- Configurable request delay, retry with exponential backoff
|
||||
- File type filtering via `--include-types` / `--exclude-types`
|
||||
- Path scoping (stay within a subdirectory) or full-site crawl
|
||||
- Subdomain crawling opt-in
|
||||
- Single-page mode (grab one page, don't crawl further)
|
||||
- Structured error logging to JSONL, kept in a separate `_logs` directory
|
||||
|
||||
## Tech Stack
|
||||
|
||||
| Thing | Choice | Why |
|
||||
|---|---|---|
|
||||
| Language | Rust 2024 | Single static binary. No runtime. No GC. No regrets. |
|
||||
| HTTP client | reqwest + tokio | Async, fast, well-maintained. |
|
||||
| HTML parsing | scraper | CSS selectors on server-side HTML. |
|
||||
| HTML→MD | htmd | Strips nav, footer, script tags. Produces clean Markdown. |
|
||||
| Document conversion | anydoc | Detects format from bytes, converts to Markdown. Falls back to raw. |
|
||||
| CLI | clap | Derive macros, typed args, good help text. |
|
||||
| Logging | env_logger | Simple, sufficient. |
|
||||
|
||||
## Usage
|
||||
|
||||
```bash
|
||||
# Scrape a single page, save Markdown
|
||||
web-scraper https://example.com/page --single
|
||||
|
||||
# Crawl entire site, only PDFs and HTML
|
||||
web-scraper https://example.com --types pdf,html
|
||||
|
||||
# Crawl a subdirectory only, exclude images
|
||||
web-scraper https://example.com/docs --exclude-types png,jpg,gif
|
||||
|
||||
# Include subdomains, custom delay
|
||||
web-scraper https://example.com --subdomains --delay-ms 500
|
||||
```
|
||||
|
||||
Output goes to `./<domain>_scraped/`, logs to `./<domain>_scraped_logs/`.
|
||||
|
||||
## CLI Flags
|
||||
|
||||
```
|
||||
START_URL Starting URL to scrape
|
||||
-o, --output DIR Output directory (default: <domain>_scraped)
|
||||
--subdomains Also crawl subdomains of the base host
|
||||
--single Only scrape the start URL, don't crawl
|
||||
--no-scope Crawl any path on the same host
|
||||
--no-doc-conversion Save documents as-is, skip conversion
|
||||
--delay-ms N Delay between requests in ms (default: 1000)
|
||||
--types a,b,c Only save these file types (HTML always crawled)
|
||||
--exclude-types a,b Skip these file types
|
||||
```
|
||||
|
||||
`--types` and `--exclude-types` are mutually exclusive. Both still crawl HTML pages for links regardless.
|
||||
|
||||
## Build
|
||||
|
||||
```bash
|
||||
just build # clippy + fmt + release build
|
||||
# or
|
||||
cargo build --release
|
||||
```
|
||||
|
||||
Release profile is optimized for size: `opt-level = "z"`, LTO, single codegen unit, stripped.
|
||||
|
||||
## Error Logs
|
||||
|
||||
Errors are logged to `scrape_errors.jsonl` in the logs directory. Each entry includes timestamp, URL, error type, message, and HTTP status code if applicable. The log file is cleared at the start of each run.
|
||||
|
||||
---
|
||||
|
||||
**Built to scrape things. Respects robots.txt. Does not require a DevOps team.**
|
||||
Reference in New Issue
Block a user