Files
hagfish/docs/TODO.md
T
2026-08-16 22:22:59 +01:00

5.3 KiB

HAGFISH Project todo

GET statbank.hagstova.fo/api/v1/fo/H2/VV/VV01/fisknv_md.px — metadata POST same URL — JSON-stat2 query 8 dimensions, "-" → NULL

Ingestion

Monthly run (cron or systemd timer, user's choice) Incremental with full backfill option GET metadata → build lookup maps → POST query → parse → insert into DuckDB Upsert semantics for revised months

Storage

DuckDB on disk

Fact table

landings(species_code, gear_code, zone_code, processing_code, preservation_code, shipsize_code, month, mass_kg, value_kr) Lookup tables: species, gear, zone, processing, preservation, shipsize Parquet export per run

API (Axum)

GET /api/species — list species codes + Faroese names GET /api/landings — filtered query, JSON response GET /api/summary — aggregates GET /api/export.parquet — download Parquet Serve embedded static frontend

Frontend

JS + ECharts

Line chart, stacked bar, donut, dropdown filters Plain HTML/CSS/JS, no build step

Deployment

Statically linked Rust binary config.json for settings (DuckDB path, bind addr, data source URL) Systemd timer for monthly ingestion (bare metal, no containers)

Project TODO

hagfish/ ├── Cargo.toml ├── Taskfile.yml ├── config.json ├── static/ │ ├── index.html │ ├── app.js │ └── style.css └── src/ ├── main.rs ├── ingest.rs ├── db.rs ├── api.rs └── types.rs

Phase 1: Types & Ingestion

  • 1.1 Define types in types.rs: MetadataResponse, VariableMeta, DataResponse, DataRow, Query, Selection, QueryItem, Config — all with serde derives
  • 1.2 Implement ingest.rs::fetch_metadata(url) — GET request, parse JSON, return HashMap<(variable_code, value_code), faroese_label>
  • 1.3 Implement ingest.rs::build_query(months: &[String]) — construct POST body with all species/gear/zones set to "*", processing/preservation/shipsize set to TOTAL, measure set to both MASS and VALUE
  • 1.4 Implement ingest.rs::fetch_data(url, query) — POST request, parse JSON-stat2 response, return Vec
  • 1.5 Implement ingest.rs::parse_row(row, lookup_maps) — decode key[] positions into labeled Landing struct. Handle "-" → None.
  • 1.6 Write unit tests: mock JSON-stat2 response, verify key-to-label mapping, verify "-" handling, verify Faroese Unicode characters in species names (ð, á, í, ý, ø, ó)

Phase 2: DuckDB Storage

  • 2.1 Add duckdb crate dependency (bundled feature)
  • 2.2 Implement db.rs::init(path) — create tables: landings fact table + 6 lookup tables. Add indexes on month, species_code.
  • 2.3 Implement db.rs::upsert_landings(rows) — batch insert with delete+insert per month or INSERT OR REPLACE
  • 2.4 Implement db.rs::update_lookups(metadata) — populate lookup tables from metadata response
  • 2.5 Implement db.rs::get_last_month() — query max month from landings table for incremental ingestion
  • 2.6 Implement db.rs::export_parquet(path) — COPY landings TO 'path' (FORMAT PARQUET) partitioned by month
  • 2.7 Write integration tests: init in-memory DB, insert sample rows, query back, verify NULL handling

Phase 3: API (Axum)

  • 3.1 Set up Axum router in main.rs with Tokio runtime. AppState holds DuckDB connection wrapped in Mutex.
  • 3.2 Implement GET /api/species — query species lookup table, return JSON array
  • 3.3 Implement GET /api/landings — parse query params, build DuckDB SQL with WHERE clauses. Support: months, species, gear, zone, measure filters.
  • 3.4 Implement GET /api/summary — aggregate query: total mass + value by month, top 10 species by value, price/kg trend
  • 3.5 Implement GET /api/export.parquet — generate and stream Parquet via DuckDB COPY
  • 3.6 Implement GET /healthz
  • 3.7 Serve static files via rust-embed
  • 3.8 Write API tests

Phase 4: Frontend

  • 4.1 index.html — dropdown filters (species, zone, gear, month range) and 3 chart containers
  • 4.2 app.js — fetch species list on load, populate dropdowns, fetch /api/landings, render charts
  • 4.3 ECharts line chart: x=month, y=mass/value toggle
  • 4.4 ECharts stacked bar: x=month, y=value by species (top 10 + "other")
  • 4.5 ECharts donut: species distribution for selected month
  • 4.6 Loading states, error handling, empty state
  • 4.7 Responsive layout, plain CSS

Phase 5: CLI & Scheduling

  • 5.1 Add clap derive subcommands: hagfish ingest [--full] and hagfish serve
  • 5.2 Implement incremental logic: read get_last_month(), compute remaining months from metadata, fetch in batches if >12 months
  • 5.3 Log ingestion runs with slog (rows inserted, duration, errors)
  • 5.4 Add hagfish export --out /path/to/parquet subcommand
  • 5.5 Load config.json on startup (DuckDB path, bind address, data source URL, log file path)

Phase 6: Bare Metal Deployment

  • 6.1 Write systemd service unit file (hagfish.service) — ExecStart=/usr/local/bin/hagfish serve, restart policy
  • 6.2 Write systemd timer (hagfish-ingest.timer + hagfish-ingest.service) — monthly, runs hagfish ingest
  • 6.3 Taskfile: build (release, static), deploy (rsync binary + config + units, ssh reload)
  • 6.4 README with ELI5 Technology Choices section (why DuckDB, why Rust, why embedded static assets)