This commit is contained in:
2026-08-16 20:53:36 +01:00
parent 53ab2c96c4
commit 24d5a4e57a
3 changed files with 793 additions and 0 deletions
+113
View File
@@ -0,0 +1,113 @@
# HAGFISH Project todo
`GET statbank.hagstova.fo/api/v1/fo/H2/VV/VV01/fisknv_md.px — metadata`
POST same URL — JSON-stat2 query
8 dimensions, "-" → NULL
## Ingestion
Monthly run (cron or systemd timer, user's choice)
Incremental with full backfill option
GET metadata → build lookup maps → POST query → parse → insert into DuckDB
Upsert semantics for revised months
## Storage
DuckDB on disk
## Fact table
landings(species_code, gear_code, zone_code, processing_code, preservation_code, shipsize_code, month, mass_kg, value_kr)
Lookup tables: species, gear, zone, processing, preservation, shipsize
Parquet export per run
## API (Axum)
GET /api/species — list species codes + Faroese names
GET /api/landings — filtered query, JSON response
GET /api/summary — aggregates
GET /api/export.parquet — download Parquet
Serve embedded static frontend
## Frontend
JS + ECharts
Line chart, stacked bar, donut, dropdown filters
Plain HTML/CSS/JS, no build step
## Deployment
Statically linked Rust binary
config.json for settings (DuckDB path, bind addr, data source URL)
Systemd timer for monthly ingestion (bare metal, no containers)
## Project TODO
hagfish/
├── Cargo.toml
├── Taskfile.yml
├── config.json
├── static/
│ ├── index.html
│ ├── app.js
│ └── style.css
└── src/
├── main.rs
├── ingest.rs
├── db.rs
├── api.rs
└── types.rs
### Phase 1: Types & Ingestion
- [ ] 1.1 Define types in types.rs: MetadataResponse, VariableMeta, DataResponse, DataRow, Query, Selection, QueryItem, Config — all with serde derives
- [ ] 1.2 Implement ingest.rs::fetch_metadata(url) — GET request, parse JSON, return HashMap<(variable_code, value_code), faroese_label>
- [ ] 1.3 Implement ingest.rs::build_query(months: &[String]) — construct POST body with all species/gear/zones set to "*", processing/preservation/shipsize set to TOTAL, measure set to both MASS and VALUE
- [ ] 1.4 Implement ingest.rs::fetch_data(url, query) — POST request, parse JSON-stat2 response, return Vec<DataRow>
- [ ] 1.5 Implement ingest.rs::parse_row(row, lookup_maps) — decode key[] positions into labeled Landing struct. Handle "-" → None.
- [ ] 1.6 Write unit tests: mock JSON-stat2 response, verify key-to-label mapping, verify "-" handling, verify Faroese Unicode characters in species names (ð, á, í, ý, ø, ó)
### Phase 2: DuckDB Storage
- [ ] 2.1 Add duckdb crate dependency (bundled feature)
- [ ] 2.2 Implement db.rs::init(path) — create tables: landings fact table + 6 lookup tables. Add indexes on month, species_code.
- [ ] 2.3 Implement db.rs::upsert_landings(rows) — batch insert with delete+insert per month or INSERT OR REPLACE
- [ ] 2.4 Implement db.rs::update_lookups(metadata) — populate lookup tables from metadata response
- [ ] 2.5 Implement db.rs::get_last_month() — query max month from landings table for incremental ingestion
- [ ] 2.6 Implement db.rs::export_parquet(path) — COPY landings TO 'path' (FORMAT PARQUET) partitioned by month
- [ ] 2.7 Write integration tests: init in-memory DB, insert sample rows, query back, verify NULL handling
### Phase 3: API (Axum)
- [ ] 3.1 Set up Axum router in main.rs with Tokio runtime. AppState holds DuckDB connection wrapped in Mutex.
- [ ] 3.2 Implement GET /api/species — query species lookup table, return JSON array
- [ ] 3.3 Implement GET /api/landings — parse query params, build DuckDB SQL with WHERE clauses. Support: months, species, gear, zone, measure filters.
- [ ] 3.4 Implement GET /api/summary — aggregate query: total mass + value by month, top 10 species by value, price/kg trend
- [ ] 3.5 Implement GET /api/export.parquet — generate and stream Parquet via DuckDB COPY
- [ ] 3.6 Implement GET /healthz
- [ ] 3.7 Serve static files via rust-embed
- [ ] 3.8 Write API tests
### Phase 4: Frontend
- [ ] 4.1 index.html — dropdown filters (species, zone, gear, month range) and 3 chart containers
- [ ] 4.2 app.js — fetch species list on load, populate dropdowns, fetch /api/landings, render charts
- [ ] 4.3 ECharts line chart: x=month, y=mass/value toggle
- [ ] 4.4 ECharts stacked bar: x=month, y=value by species (top 10 + "other")
- [ ] 4.5 ECharts donut: species distribution for selected month
- [ ] 4.6 Loading states, error handling, empty state
- [ ] 4.7 Responsive layout, plain CSS
### Phase 5: CLI & Scheduling
- [ ] 5.1 Add clap derive subcommands: hagfish ingest [--full] and hagfish serve
- [ ] 5.2 Implement incremental logic: read get_last_month(), compute remaining months from metadata, fetch in batches if >12 months
- [ ] 5.3 Log ingestion runs with slog (rows inserted, duration, errors)
- [ ] 5.4 Add hagfish export --out /path/to/parquet subcommand
- [ ] 5.5 Load config.json on startup (DuckDB path, bind address, data source URL, log file path)
### Phase 6: Bare Metal Deployment
- [ ] 6.1 Write systemd service unit file (hagfish.service) — ExecStart=/usr/local/bin/hagfish serve, restart policy
- [ ] 6.2 Write systemd timer (hagfish-ingest.timer + hagfish-ingest.service) — monthly, runs hagfish ingest
- [ ] 6.3 Taskfile: build (release, static), deploy (rsync binary + config + units, ssh reload)
- [ ] 6.4 README with ELI5 Technology Choices section (why DuckDB, why Rust, why embedded static assets)