readem
This commit is contained in:
@@ -0,0 +1,83 @@
|
|||||||
|
# web-scraper-rust
|
||||||
|
|
||||||
|
A polite, single-binary Rust web scraper that crawls pages, converts HTML to Markdown, and tries its best to turn Office documents into something readable too.
|
||||||
|
|
||||||
|
It respects `robots.txt`. It waits between requests. It hashes long URLs into filenames because filesystems have feelings too. It does not do cookies, JavaScript rendering, or anything fancy. If you need a headless browser, this is not your tool.
|
||||||
|
|
||||||
|
## What It Does
|
||||||
|
|
||||||
|
- Crawls internal links from a starting URL (BFS queue)
|
||||||
|
- Converts HTML pages to Markdown via `htmd`
|
||||||
|
- Detects and converts documents (PDF, DOCX, XLSX, PPTX, ODT, RTF, EPUB, CSV...) to Markdown via `anydoc`
|
||||||
|
- Respects `robots.txt` with wildcard pattern matching
|
||||||
|
- Deduplicates URLs, normalizes fragments and trailing query params
|
||||||
|
- Configurable request delay, retry with exponential backoff
|
||||||
|
- File type filtering via `--include-types` / `--exclude-types`
|
||||||
|
- Path scoping (stay within a subdirectory) or full-site crawl
|
||||||
|
- Subdomain crawling opt-in
|
||||||
|
- Single-page mode (grab one page, don't crawl further)
|
||||||
|
- Structured error logging to JSONL, kept in a separate `_logs` directory
|
||||||
|
|
||||||
|
## Tech Stack
|
||||||
|
|
||||||
|
| Thing | Choice | Why |
|
||||||
|
|---|---|---|
|
||||||
|
| Language | Rust 2024 | Single static binary. No runtime. No GC. No regrets. |
|
||||||
|
| HTTP client | reqwest + tokio | Async, fast, well-maintained. |
|
||||||
|
| HTML parsing | scraper | CSS selectors on server-side HTML. |
|
||||||
|
| HTML→MD | htmd | Strips nav, footer, script tags. Produces clean Markdown. |
|
||||||
|
| Document conversion | anydoc | Detects format from bytes, converts to Markdown. Falls back to raw. |
|
||||||
|
| CLI | clap | Derive macros, typed args, good help text. |
|
||||||
|
| Logging | env_logger | Simple, sufficient. |
|
||||||
|
|
||||||
|
## Usage
|
||||||
|
|
||||||
|
```bash
|
||||||
|
# Scrape a single page, save Markdown
|
||||||
|
web-scraper https://example.com/page --single
|
||||||
|
|
||||||
|
# Crawl entire site, only PDFs and HTML
|
||||||
|
web-scraper https://example.com --types pdf,html
|
||||||
|
|
||||||
|
# Crawl a subdirectory only, exclude images
|
||||||
|
web-scraper https://example.com/docs --exclude-types png,jpg,gif
|
||||||
|
|
||||||
|
# Include subdomains, custom delay
|
||||||
|
web-scraper https://example.com --subdomains --delay-ms 500
|
||||||
|
```
|
||||||
|
|
||||||
|
Output goes to `./<domain>_scraped/`, logs to `./<domain>_scraped_logs/`.
|
||||||
|
|
||||||
|
## CLI Flags
|
||||||
|
|
||||||
|
```
|
||||||
|
START_URL Starting URL to scrape
|
||||||
|
-o, --output DIR Output directory (default: <domain>_scraped)
|
||||||
|
--subdomains Also crawl subdomains of the base host
|
||||||
|
--single Only scrape the start URL, don't crawl
|
||||||
|
--no-scope Crawl any path on the same host
|
||||||
|
--no-doc-conversion Save documents as-is, skip conversion
|
||||||
|
--delay-ms N Delay between requests in ms (default: 1000)
|
||||||
|
--types a,b,c Only save these file types (HTML always crawled)
|
||||||
|
--exclude-types a,b Skip these file types
|
||||||
|
```
|
||||||
|
|
||||||
|
`--types` and `--exclude-types` are mutually exclusive. Both still crawl HTML pages for links regardless.
|
||||||
|
|
||||||
|
## Build
|
||||||
|
|
||||||
|
```bash
|
||||||
|
just build # clippy + fmt + release build
|
||||||
|
# or
|
||||||
|
cargo build --release
|
||||||
|
```
|
||||||
|
|
||||||
|
Release profile is optimized for size: `opt-level = "z"`, LTO, single codegen unit, stripped.
|
||||||
|
|
||||||
|
## Error Logs
|
||||||
|
|
||||||
|
Errors are logged to `scrape_errors.jsonl` in the logs directory. Each entry includes timestamp, URL, error type, message, and HTTP status code if applicable. The log file is cleared at the start of each run.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
**Built to scrape things. Respects robots.txt. Does not require a DevOps team.**
|
||||||
Reference in New Issue
Block a user