# HAGFISH Project todo `GET statbank.hagstova.fo/api/v1/fo/H2/VV/VV01/fisknv_md.px — metadata` POST same URL — JSON-stat2 query 8 dimensions, "-" → NULL ## Ingestion Monthly run (cron or systemd timer, user's choice) Incremental with full backfill option GET metadata → build lookup maps → POST query → parse → insert into DuckDB Upsert semantics for revised months ## Storage DuckDB on disk ## Fact table landings(species_code, gear_code, zone_code, processing_code, preservation_code, shipsize_code, month, mass_kg, value_kr) Lookup tables: species, gear, zone, processing, preservation, shipsize Parquet export per run ## API (Axum) GET /api/species — list species codes + Faroese names GET /api/landings — filtered query, JSON response GET /api/summary — aggregates GET /api/export.parquet — download Parquet Serve embedded static frontend ## Frontend JS + ECharts Line chart, stacked bar, donut, dropdown filters Plain HTML/CSS/JS, no build step ## Deployment Statically linked Rust binary config.json for settings (DuckDB path, bind addr, data source URL) Systemd timer for monthly ingestion (bare metal, no containers) ## Project TODO hagfish/ ├── Cargo.toml ├── Taskfile.yml ├── config.json ├── static/ │ ├── index.html │ ├── app.js │ └── style.css └── src/ ├── main.rs ├── ingest.rs ├── db.rs ├── api.rs └── types.rs ### Phase 1: Types & Ingestion - [x] 1.1 Define types in types.rs: MetadataResponse, VariableMeta, DataResponse, DataRow, Query, Selection, QueryItem, Config — all with serde derives - [x] 1.2 Implement ingest.rs::fetch_metadata(url) — GET request, parse JSON, return HashMap<(variable_code, value_code), faroese_label> - [x] 1.3 Implement ingest.rs::build_query(months: &[String]) — construct POST body with all species/gear/zones set to "*", processing/preservation/shipsize set to TOTAL, measure set to both MASS and VALUE - [x] 1.4 Implement ingest.rs::fetch_data(url, query) — POST request, parse JSON-stat2 response, return Vec - [x] 1.5 Implement ingest.rs::parse_row(row, lookup_maps) — decode key[] positions into labeled Landing struct. Handle "-" → None. - [x] 1.6 Write unit tests: mock JSON-stat2 response, verify key-to-label mapping, verify "-" handling, verify Faroese Unicode characters in species names (ð, á, í, ý, ø, ó) ### Phase 2: DuckDB Storage - [x] 2.1 Add duckdb crate dependency (bundled feature) - [x] 2.2 Implement db.rs::init(path) — create tables: landings fact table + 6 lookup tables. Add indexes on month, species_code. - [x] 2.3 Implement db.rs::upsert_landings(rows) — batch insert with delete+insert per month or INSERT OR REPLACE - [x] 2.4 Implement db.rs::update_lookups(metadata) — populate lookup tables from metadata response - [x] 2.5 Implement db.rs::get_last_month() — query max month from landings table for incremental ingestion - [x] 2.6 Implement db.rs::export_parquet(path) — COPY landings TO 'path' (FORMAT PARQUET) partitioned by month - [x] 2.7 Write integration tests: init in-memory DB, insert sample rows, query back, verify NULL handling ### Phase 3: API (Axum) - [ ] 3.1 Set up Axum router in main.rs with Tokio runtime. AppState holds DuckDB connection wrapped in Mutex. - [ ] 3.2 Implement GET /api/species — query species lookup table, return JSON array - [ ] 3.3 Implement GET /api/landings — parse query params, build DuckDB SQL with WHERE clauses. Support: months, species, gear, zone, measure filters. - [ ] 3.4 Implement GET /api/summary — aggregate query: total mass + value by month, top 10 species by value, price/kg trend - [ ] 3.5 Implement GET /api/export.parquet — generate and stream Parquet via DuckDB COPY - [ ] 3.6 Implement GET /healthz - [ ] 3.7 Serve static files via rust-embed - [ ] 3.8 Write API tests ### Phase 4: Frontend - [ ] 4.1 index.html — dropdown filters (species, zone, gear, month range) and 3 chart containers - [ ] 4.2 app.js — fetch species list on load, populate dropdowns, fetch /api/landings, render charts - [ ] 4.3 ECharts line chart: x=month, y=mass/value toggle - [ ] 4.4 ECharts stacked bar: x=month, y=value by species (top 10 + "other") - [ ] 4.5 ECharts donut: species distribution for selected month - [ ] 4.6 Loading states, error handling, empty state - [ ] 4.7 Responsive layout, plain CSS ### Phase 5: CLI & Scheduling - [ ] 5.1 Add clap derive subcommands: hagfish ingest [--full] and hagfish serve - [ ] 5.2 Implement incremental logic: read get_last_month(), compute remaining months from metadata, fetch in batches if >12 months - [ ] 5.3 Log ingestion runs with slog (rows inserted, duration, errors) - [ ] 5.4 Add hagfish export --out /path/to/parquet subcommand - [ ] 5.5 Load config.json on startup (DuckDB path, bind address, data source URL, log file path) ### Phase 6: Bare Metal Deployment - [ ] 6.1 Write systemd service unit file (hagfish.service) — ExecStart=/usr/local/bin/hagfish serve, restart policy - [ ] 6.2 Write systemd timer (hagfish-ingest.timer + hagfish-ingest.service) — monthly, runs hagfish ingest - [ ] 6.3 Taskfile: build (release, static), deploy (rsync binary + config + units, ssh reload) - [ ] 6.4 README with ELI5 Technology Choices section (why DuckDB, why Rust, why embedded static assets)