114 lines
5.3 KiB
Markdown
114 lines
5.3 KiB
Markdown
# HAGFISH Project todo
|
|
|
|
`GET statbank.hagstova.fo/api/v1/fo/H2/VV/VV01/fisknv_md.px — metadata`
|
|
POST same URL — JSON-stat2 query
|
|
8 dimensions, "-" → NULL
|
|
|
|
## Ingestion
|
|
|
|
Monthly run (cron or systemd timer, user's choice)
|
|
Incremental with full backfill option
|
|
GET metadata → build lookup maps → POST query → parse → insert into DuckDB
|
|
Upsert semantics for revised months
|
|
|
|
## Storage
|
|
DuckDB on disk
|
|
|
|
## Fact table
|
|
landings(species_code, gear_code, zone_code, processing_code, preservation_code, shipsize_code, month, mass_kg, value_kr)
|
|
Lookup tables: species, gear, zone, processing, preservation, shipsize
|
|
Parquet export per run
|
|
|
|
## API (Axum)
|
|
|
|
GET /api/species — list species codes + Faroese names
|
|
GET /api/landings — filtered query, JSON response
|
|
GET /api/summary — aggregates
|
|
GET /api/export.parquet — download Parquet
|
|
Serve embedded static frontend
|
|
|
|
## Frontend
|
|
JS + ECharts
|
|
|
|
Line chart, stacked bar, donut, dropdown filters
|
|
Plain HTML/CSS/JS, no build step
|
|
|
|
## Deployment
|
|
|
|
Statically linked Rust binary
|
|
config.json for settings (DuckDB path, bind addr, data source URL)
|
|
Systemd timer for monthly ingestion (bare metal, no containers)
|
|
|
|
|
|
## Project TODO
|
|
|
|
hagfish/
|
|
├── Cargo.toml
|
|
├── Taskfile.yml
|
|
├── config.json
|
|
├── static/
|
|
│ ├── index.html
|
|
│ ├── app.js
|
|
│ └── style.css
|
|
└── src/
|
|
├── main.rs
|
|
├── ingest.rs
|
|
├── db.rs
|
|
├── api.rs
|
|
└── types.rs
|
|
|
|
### Phase 1: Types & Ingestion
|
|
|
|
- [ ] 1.1 Define types in types.rs: MetadataResponse, VariableMeta, DataResponse, DataRow, Query, Selection, QueryItem, Config — all with serde derives
|
|
- [ ] 1.2 Implement ingest.rs::fetch_metadata(url) — GET request, parse JSON, return HashMap<(variable_code, value_code), faroese_label>
|
|
- [ ] 1.3 Implement ingest.rs::build_query(months: &[String]) — construct POST body with all species/gear/zones set to "*", processing/preservation/shipsize set to TOTAL, measure set to both MASS and VALUE
|
|
- [ ] 1.4 Implement ingest.rs::fetch_data(url, query) — POST request, parse JSON-stat2 response, return Vec<DataRow>
|
|
- [ ] 1.5 Implement ingest.rs::parse_row(row, lookup_maps) — decode key[] positions into labeled Landing struct. Handle "-" → None.
|
|
- [ ] 1.6 Write unit tests: mock JSON-stat2 response, verify key-to-label mapping, verify "-" handling, verify Faroese Unicode characters in species names (ð, á, í, ý, ø, ó)
|
|
|
|
### Phase 2: DuckDB Storage
|
|
|
|
- [ ] 2.1 Add duckdb crate dependency (bundled feature)
|
|
- [ ] 2.2 Implement db.rs::init(path) — create tables: landings fact table + 6 lookup tables. Add indexes on month, species_code.
|
|
- [ ] 2.3 Implement db.rs::upsert_landings(rows) — batch insert with delete+insert per month or INSERT OR REPLACE
|
|
- [ ] 2.4 Implement db.rs::update_lookups(metadata) — populate lookup tables from metadata response
|
|
- [ ] 2.5 Implement db.rs::get_last_month() — query max month from landings table for incremental ingestion
|
|
- [ ] 2.6 Implement db.rs::export_parquet(path) — COPY landings TO 'path' (FORMAT PARQUET) partitioned by month
|
|
- [ ] 2.7 Write integration tests: init in-memory DB, insert sample rows, query back, verify NULL handling
|
|
|
|
### Phase 3: API (Axum)
|
|
|
|
- [ ] 3.1 Set up Axum router in main.rs with Tokio runtime. AppState holds DuckDB connection wrapped in Mutex.
|
|
- [ ] 3.2 Implement GET /api/species — query species lookup table, return JSON array
|
|
- [ ] 3.3 Implement GET /api/landings — parse query params, build DuckDB SQL with WHERE clauses. Support: months, species, gear, zone, measure filters.
|
|
- [ ] 3.4 Implement GET /api/summary — aggregate query: total mass + value by month, top 10 species by value, price/kg trend
|
|
- [ ] 3.5 Implement GET /api/export.parquet — generate and stream Parquet via DuckDB COPY
|
|
- [ ] 3.6 Implement GET /healthz
|
|
- [ ] 3.7 Serve static files via rust-embed
|
|
- [ ] 3.8 Write API tests
|
|
|
|
### Phase 4: Frontend
|
|
|
|
- [ ] 4.1 index.html — dropdown filters (species, zone, gear, month range) and 3 chart containers
|
|
- [ ] 4.2 app.js — fetch species list on load, populate dropdowns, fetch /api/landings, render charts
|
|
- [ ] 4.3 ECharts line chart: x=month, y=mass/value toggle
|
|
- [ ] 4.4 ECharts stacked bar: x=month, y=value by species (top 10 + "other")
|
|
- [ ] 4.5 ECharts donut: species distribution for selected month
|
|
- [ ] 4.6 Loading states, error handling, empty state
|
|
- [ ] 4.7 Responsive layout, plain CSS
|
|
|
|
### Phase 5: CLI & Scheduling
|
|
|
|
- [ ] 5.1 Add clap derive subcommands: hagfish ingest [--full] and hagfish serve
|
|
- [ ] 5.2 Implement incremental logic: read get_last_month(), compute remaining months from metadata, fetch in batches if >12 months
|
|
- [ ] 5.3 Log ingestion runs with slog (rows inserted, duration, errors)
|
|
- [ ] 5.4 Add hagfish export --out /path/to/parquet subcommand
|
|
- [ ] 5.5 Load config.json on startup (DuckDB path, bind address, data source URL, log file path)
|
|
|
|
### Phase 6: Bare Metal Deployment
|
|
|
|
- [ ] 6.1 Write systemd service unit file (hagfish.service) — ExecStart=/usr/local/bin/hagfish serve, restart policy
|
|
- [ ] 6.2 Write systemd timer (hagfish-ingest.timer + hagfish-ingest.service) — monthly, runs hagfish ingest
|
|
- [ ] 6.3 Taskfile: build (release, static), deploy (rsync binary + config + units, ssh reload)
|
|
- [ ] 6.4 README with ELI5 Technology Choices section (why DuckDB, why Rust, why embedded static assets)
|