phase 1 complete

This commit is contained in:
2026-08-16 22:22:59 +01:00
parent 24d5a4e57a
commit 5ec60a71c3
6 changed files with 2707 additions and 341 deletions
Generated
+1838
View File
File diff suppressed because it is too large Load Diff
+192
View File
@@ -0,0 +1,192 @@
# API Documentation
This document describes the PX-Web API used by hagfish to fetch Faroese fisheries statistics from the official Statbank.
## EndpointBase URL: https://statbank.hagstova.fo/api/v1/fo/H2/VV/VV01/fisknv_md.px
- **GET**: Returns metadata (table structure, dimension codes, labels)
- **POST**: Returns data in JSON-stat2 format
## Metadata Request (GET)
### Requestbash
curl -s "https://statbank.hagstova.fo/api/v1/fo/H2/VV/VV01/fisknv_md.px"
### Response Structurejson
{
"title": "AVR01010 Avreiðingar í nøgd og virði...",
"variables": [
{
"code": "measure",
"text": "mát",
"role": null,
"values": ["MASS", "VALUE"],
"valueTexts": ["Nøgd", "Virði"]
}
]
}
### Fields
| Field | Type | Description |
|-------|------|-------------|
| `title` | string | Table name in Faroese |
| `variables[].code` | string | Dimension identifier (used in queries) |
| `variables[].text` | string | Human-readable label in Faroese |
| `variables[].role` | null/string | Classification (always `null` on this endpoint) |
| `variables[].values` | string[] | Valid dimension codes |
| `variables[].valueTexts` | string[] | Display labels (parallel with `values`) |
### Verified Dimension Codes
| Variable Code | Faroese Label | Values |
|---------------|---------------|--------|
| `measure` | mát | `MASS`, `VALUE` |
| `Species (ASFIS2022)` | Fiskaslag (ASFIS2022) | `TOTAL`, `148XXXXXXX00000`, ... (72 total) |
| `Fishing Gear (ISSCFG2016)` | Reiðskapur (ISSCFG2016) | `TOTAL`, ... (12 total) |
| `Economic Zone (GEONOM2023)` | Búskapar øki (GEONOM2023) | `TOTAL`, ... (10 total) |
| `Processing (EUMOFAPresentation)` | Virking (EUMOFAPresentation) | `TOTAL`, ... (6 total) |
| `Preservation (EUMOFAPreservation)` | Viðgerð (EUMOFAPreservation) | `TOTAL`, ... (9 total) |
| `Shipsize` | Skipastødd | `TOTAL`, ... (11 total) |
| `month` | mánaður | `2015M01` through `2026M05` (137 months) |
**Note**: `role` is always `null`. Time dimension identification must use `code == "month"`.
---
## Data Request (POST)
### Request Body Formatjson
{
"query": [
{
"code": "dimension_code",
"selection": {
"filter": "item",
"values": ["value1", "value2"]
}
}
],
"response": {
"format": "json-stat2"
}
}
### Filter Types
| Filter | Values | Description |
|--------|--------|-------------|
| `item` | explicit codes | Select specific values |
| `all` | `["*"]` | Wildcard — select all values |
| `top` | `["TOP_N"]` | Top N values by magnitude |
### Example Querybash
curl -s -X POST
-H "Content-Type: application/json"
-d '{
"query": [
{"code": "month", "selection": {"filter": "item", "values": ["2024M01", "2024M02"]}},
{"code": "Species (ASFIS2022)", "selection": {"filter": "all", "values": [""]}},
{"code": "Fishing Gear (ISSCFG2016)", "selection": {"filter": "all", "values": [""]}},
{"code": "Economic Zone (GEONOM2023)", "selection": {"filter": "all", "values": ["*"]}},
{"code": "Processing (EUMOFAPresentation)", "selection": {"filter": "item", "values": ["TOTAL"]}},
{"code": "Preservation (EUMOFAPreservation)", "selection": {"filter": "item", "values": ["TOTAL"]}},
{"code": "Shipsize", "selection": {"filter": "item", "values": ["TOTAL"]}},
{"code": "measure", "selection": {"filter": "item", "values": ["MASS", "VALUE"]}}
],
"response": {"format": "json-stat2"}
}'
"https://statbank.hagstova.fo/api/v1/fo/H2/VV/VV01/fisknv_md.px"
### Cell Limit
- Maximum cells per query: ~8,000,000
- Recommended maximum: 1,000,000 for reliability
- To fetch full dataset, paginate by month or use smaller dimension selections
---
## Response Format (JSON-stat2)
### Structurejson
{
"class": "dataset",
"label": "...",
"id": ["measure", "Species (ASFIS2022)", ..., "month"],
"size": [2, 72, 12, 10, 1, 1, 1, 2],
"dimension": {
"measure": {
"label": "mát",
"category": {
"index": {"MASS": 0, "VALUE": 1},
"label": {"MASS": "Nøgd", "VALUE": "Virði"}
}
}
},
"value": [73653738, 1234567, ...]
}
### Fields
| Field | Description |
|-------|-------------|
| `id` | Array of dimension codes in order (defines cube layout) |
| `size` | Cardinality per dimension (parallel with `id`) |
| `dimension.<code>.category.index` | Maps value code → numeric position in dimension |
| `dimension.<code>.category.label` | Maps value code → display text |
| `value` | Flattened array of measurements (row-major order) |
### Row-Major Index Decoding
Values are stored in row-major order: the last dimension (`month`) varies fastest.
Given `size = [2, 72, 12, 10, 1, 1, 1, 2]`:
- Flat index `0` → `[0,0,0,0,0,0,0,0]` = measure=MASS, species=TOTAL, ..., month=2024M01
- Flat index `1` → `[0,0,0,0,0,0,0,1]` = measure=MASS, species=TOTAL, ..., month=2024M02
- Flat index `2` → `[0,0,0,0,0,0,1,0]` = measure=MASS, species=SPECIES_1, ..., month=2024M01
Decoding algorithm (reverse modulo):rust
fn decode(flat_index: usize, sizes: &[usize]) -> Vec<usize> {
let mut indices = Vec::new();
let mut remaining = flat_index;
for &size in sizes.iter().rev() {
indices.push(remaining % size);
remaining /= size;
}
indices.reverse();
indices
}
### Sentinel Values
| Value | Meaning |
|-------|---------|
| `-1.0` | Missing/no data (treated as `NULL`) |
| `null` | Missing/no data (already nullable in JSON) |
---
## Known Constraints
1. **Cell limit**: ~8 million cells per query
2. **Rate limits**: Unknown — assume reasonable backoff for large downloads
3. **Language**: API responds in requested language (`fo` or `en`)
4. **Time zone**: No timezone specified — treat timestamps as local
5. **Updates**: Data updated irregularly — metadata timestamp in response header
---
## Integration Checklist
- [x] Metadata endpoint reachable via GET
- [x] POST queries return valid JSON-stat2
- [x] Dimension codes match `variables[].code` from metadata
- [x] Row-major index decoding verified
- [x] Sentinel value coercion (`-1.0` → `None`) implemented
- [x] Unicode (Faroese characters) handled correctly
- [ ] Incremental ingestion logic tested
- [ ] Parquet export tested
- [ ] Error handling for network timeouts implemented
---
## References
- Official PxWeb documentation: https://pxweb.github.io/docs/
- JSON-stat2 specification: http://json-stat.org/format/
- Hagstova Føroya: https://www.hagstova.fo/
+6 -6
View File
@@ -59,12 +59,12 @@ hagfish/
### Phase 1: Types & Ingestion ### Phase 1: Types & Ingestion
- [ ] 1.1 Define types in types.rs: MetadataResponse, VariableMeta, DataResponse, DataRow, Query, Selection, QueryItem, Config — all with serde derives - [x] 1.1 Define types in types.rs: MetadataResponse, VariableMeta, DataResponse, DataRow, Query, Selection, QueryItem, Config — all with serde derives
- [ ] 1.2 Implement ingest.rs::fetch_metadata(url) — GET request, parse JSON, return HashMap<(variable_code, value_code), faroese_label> - [x] 1.2 Implement ingest.rs::fetch_metadata(url) — GET request, parse JSON, return HashMap<(variable_code, value_code), faroese_label>
- [ ] 1.3 Implement ingest.rs::build_query(months: &[String]) — construct POST body with all species/gear/zones set to "*", processing/preservation/shipsize set to TOTAL, measure set to both MASS and VALUE - [x] 1.3 Implement ingest.rs::build_query(months: &[String]) — construct POST body with all species/gear/zones set to "*", processing/preservation/shipsize set to TOTAL, measure set to both MASS and VALUE
- [ ] 1.4 Implement ingest.rs::fetch_data(url, query) — POST request, parse JSON-stat2 response, return Vec<DataRow> - [x] 1.4 Implement ingest.rs::fetch_data(url, query) — POST request, parse JSON-stat2 response, return Vec<DataRow>
- [ ] 1.5 Implement ingest.rs::parse_row(row, lookup_maps) — decode key[] positions into labeled Landing struct. Handle "-" → None. - [x] 1.5 Implement ingest.rs::parse_row(row, lookup_maps) — decode key[] positions into labeled Landing struct. Handle "-" → None.
- [ ] 1.6 Write unit tests: mock JSON-stat2 response, verify key-to-label mapping, verify "-" handling, verify Faroese Unicode characters in species names (ð, á, í, ý, ø, ó) - [x] 1.6 Write unit tests: mock JSON-stat2 response, verify key-to-label mapping, verify "-" handling, verify Faroese Unicode characters in species names (ð, á, í, ý, ø, ó)
### Phase 2: DuckDB Storage ### Phase 2: DuckDB Storage
+483 -251
View File
@@ -2,18 +2,49 @@
//! //!
//! Handles metadata fetching, query construction, data retrieval, and //! Handles metadata fetching, query construction, data retrieval, and
//! parsing of JSON-stat2 responses into structured rows. //! parsing of JSON-stat2 responses into structured rows.
//!
//! ## API contract (confirmed 2026-08-16)
//!
//! - GET endpoint: `https://statbank.hagstova.fo/api/v1/fo/H2/VV/VV01/fisknv_md.px`
//! - POST endpoint: same URL
//! - Dimension codes are NOT Faroese short names — they use classification
//! identifiers like "Species (ASFIS2022)", "Fishing Gear (ISSCFG2016)".
//! - `role` is null for all variables on this endpoint.
//! - Metadata uses parallel `values`/`valueTexts` string arrays.
//! - Sentinel values: -1.0 indicates missing data.
use crate::types::*; use crate::types::*;
use reqwest::Client; use reqwest::Client;
use tracing::{debug, info, warn};
use std::collections::HashMap; use std::collections::HashMap;
use tracing::{debug, info, warn};
const PX_WEB_LANGUAGE: &str = "fo"; // Faroese labels /// Language code for PX-Web queries.
const PX_WEB_LANGUAGE: &str = "fo";
/// Dimension codes — confirmed against live API on 2026-08-16.
/// These are the `code` field values from the metadata response.
const DIM_MONTH: &str = "month";
const DIM_SPECIES: &str = "Species (ASFIS2022)";
const DIM_GEAR: &str = "Fishing Gear (ISSCFG2016)";
const DIM_ZONE: &str = "Economic Zone (GEONOM2023)";
const DIM_PROCESSING: &str = "Processing (EUMOFAPresentation)";
const DIM_PRESERVATION: &str = "Preservation (EUMOFAPreservation)";
const DIM_SHIPSIZE: &str = "Shipsize";
const DIM_MEASURE: &str = "measure";
/// Sentinel f64 values that indicate missing data, coerced to None.
const SENTINEL_VALUES: [f64; 2] = [-1.0, f64::NAN];
/// Checks whether a numeric value is a sentinel.
fn is_sentinel(v: f64) -> bool {
SENTINEL_VALUES
.iter()
.any(|s| v.total_cmp(s) == std::cmp::Ordering::Equal)
}
/// Fetches metadata from PX-Web API endpoint. /// Fetches metadata from PX-Web API endpoint.
/// ///
/// Returns a HashMap mapping variable_id → code → label mappings. /// Returns a `LookupMap` mapping dimension code → (value code → label).
/// This lookup is used to decode opaque codes in data responses.
pub async fn fetch_metadata(client: &Client, url: &str) -> Result<LookupMap> { pub async fn fetch_metadata(client: &Client, url: &str) -> Result<LookupMap> {
info!("Fetching metadata from {}", url); info!("Fetching metadata from {}", url);
@@ -24,10 +55,11 @@ pub async fn fetch_metadata(client: &Client, url: &str) -> Result<LookupMap> {
.await?; .await?;
if !resp.status().is_success() { if !resp.status().is_success() {
let body = resp.text().await.unwrap_or_default();
return Err(IngestError::HttpError(reqwest::Error::from( return Err(IngestError::HttpError(reqwest::Error::from(
std::io::Error::new( std::io::Error::new(
std::io::ErrorKind::Other, std::io::ErrorKind::Other,
format!("API returned status: {}", resp.status()), format!("API returned error: {}", body),
), ),
))); )));
} }
@@ -35,79 +67,131 @@ pub async fn fetch_metadata(client: &Client, url: &str) -> Result<LookupMap> {
let meta: MetadataResponse = resp.json().await?; let meta: MetadataResponse = resp.json().await?;
debug!("Received {} variables from metadata", meta.variables.len()); debug!("Received {} variables from metadata", meta.variables.len());
// Build lookup map from metadata for var in &meta.variables {
let mut lookup_map: LookupMap = HashMap::new(); info!(
"Variable: code='{}' text='{}' role={:?} values={}",
var.code,
var.text,
var.role,
var.values.len()
);
}
let mut lookup_map: LookupMap = HashMap::with_capacity(meta.variables.len());
for var in &meta.variables { for var in &meta.variables {
let codes = meta.values.get(&var.id); if var.values.is_empty() {
if let Some(codes) = codes { debug!("Variable '{}' has no values", var.code);
let labels: HashMap<String, String> = codes continue;
.iter()
.map(|vm| (vm.code.clone(), vm.text.clone()))
.collect();
lookup_map.insert(var.id.clone(), labels);
info!("Loaded {} codes for variable '{}'", labels.len(), var.id);
} else {
warn!("No codes found for variable '{}'", var.id);
} }
let lookup = var.lookup_map();
info!("Loaded {} codes for '{}'", lookup.len(), var.code);
lookup_map.insert(var.code.clone(), lookup);
} }
Ok(lookup_map) Ok(lookup_map)
} }
/// Extracts all available months from metadata.
///
/// Uses the "month" dimension code since `role` is null on this endpoint.
pub fn extract_available_months(meta: &MetadataResponse) -> Vec<String> {
meta.variables
.iter()
.find(|v| v.code == DIM_MONTH)
.map(|v| v.values.clone())
.unwrap_or_default()
}
/// Constructs a query body for fetching all landing data. /// Constructs a query body for fetching all landing data.
/// ///
/// Sets all categorical dimensions to "*" (all values) and measures to MASS+VALUE. /// Sets categorical dimensions to wildcard ("all" filter), time to explicit
/// Used for full backfill ingestion. /// months, and measure to MASS + VALUE.
///
/// # Panics
/// Panics if `all_months` is empty.
pub fn build_query(all_months: &[String]) -> Query { pub fn build_query(all_months: &[String]) -> Query {
assert!(
!all_months.is_empty(),
"build_query requires at least one month"
);
Query { Query {
query: vec![ query: vec![
QueryItem { QueryItem {
id: "Tid".to_string(), // Month code: DIM_MONTH.to_string(),
selection: Selection {
filter: "item".to_string(),
values: all_months.to_vec(), values: all_months.to_vec(),
}, },
},
QueryItem { QueryItem {
id: "Art".to_string(), // Species code: DIM_SPECIES.to_string(),
selection: Selection {
filter: "all".to_string(),
values: vec!["*".to_string()], values: vec!["*".to_string()],
}, },
},
QueryItem { QueryItem {
id: "Redskab".to_string(), // Gear code: DIM_GEAR.to_string(),
selection: Selection {
filter: "all".to_string(),
values: vec!["*".to_string()], values: vec!["*".to_string()],
}, },
},
QueryItem { QueryItem {
id: "Økonomisk zone".to_string(), code: DIM_ZONE.to_string(),
selection: Selection {
filter: "all".to_string(),
values: vec!["*".to_string()], values: vec!["*".to_string()],
}, },
QueryItem {
id: "Tilstand".to_string(), // Processing
values: vec!["TOTAL".to_string()],
}, },
QueryItem { QueryItem {
id: "Konservering".to_string(), // Preservation code: DIM_PROCESSING.to_string(),
values: vec!["TOTAL".to_string()], selection: Selection {
filter: "all".to_string(),
values: vec!["*".to_string()],
},
}, },
QueryItem { QueryItem {
id: "Skibsstørrelse".to_string(), // Vessel size code: DIM_PRESERVATION.to_string(),
values: vec!["TOTAL".to_string()], selection: Selection {
filter: "all".to_string(),
values: vec!["*".to_string()],
},
}, },
QueryItem { QueryItem {
id: "Måleenhed".to_string(), // Measure code: DIM_SHIPSIZE.to_string(),
selection: Selection {
filter: "all".to_string(),
values: vec!["*".to_string()],
},
},
QueryItem {
code: DIM_MEASURE.to_string(),
selection: Selection {
filter: "item".to_string(),
values: vec!["MASS".to_string(), "VALUE".to_string()], values: vec!["MASS".to_string(), "VALUE".to_string()],
}, },
},
], ],
language: PX_WEB_LANGUAGE.to_string(), response: QueryResponse::default(),
} }
} }
/// Fetches data from PX-Web API using the provided query. /// Fetches data from PX-Web API using the provided query.
pub async fn fetch_data(client: &Client, url: &str, query: &Query) -> Result<DataResponse> { pub async fn fetch_data(client: &Client, url: &str, query: &Query) -> Result<DataResponse> {
info!("Sending data query for {} month(s)", query.query[0].values.len()); let month_count = query
.query
.iter()
.find(|q| q.code == DIM_MONTH)
.map(|q| q.selection.values.len())
.unwrap_or(0);
let resp = client info!("Sending data query for {} month(s)", month_count);
.post(url)
.json(query) let resp = client.post(url).json(query).send().await?;
.send()
.await?;
if !resp.status().is_success() { if !resp.status().is_success() {
let body = resp.text().await.unwrap_or_default(); let body = resp.text().await.unwrap_or_default();
@@ -115,7 +199,7 @@ pub async fn fetch_data(client: &Client, url: &str, query: &Query) -> Result<Dat
return Err(IngestError::HttpError(reqwest::Error::from( return Err(IngestError::HttpError(reqwest::Error::from(
std::io::Error::new( std::io::Error::new(
std::io::ErrorKind::Other, std::io::ErrorKind::Other,
format!("API returned status: {}", resp.status()), format!("API returned error: {}", body),
), ),
))); )));
} }
@@ -126,53 +210,51 @@ pub async fn fetch_data(client: &Client, url: &str, query: &Query) -> Result<Dat
return Err(IngestError::EmptyDataset); return Err(IngestError::EmptyDataset);
} }
info!( info!("Retrieved {} data points", data.dataset.value.len());
"Retrieved {} data points from API",
data.dataset.value.len()
);
Ok(data) Ok(data)
} }
/// Extracts dimension codes from a flat row index in JSON-stat2 format. /// Decodes a flat row-major index into per-dimension indices.
/// ///
/// JSON-stat2 uses a flattened array where each element corresponds to a unique /// JSON-stat2 uses row-major order: the last dimension varies fastest.
/// combination of dimension values. The dimension.keys array tells us how many fn decode_key_indices(flat_index: usize, key_sizes: &[usize]) -> Vec<usize> {
/// values exist per dimension. We decode the flat index into per-dimension indices. let mut indices = Vec::with_capacity(key_sizes.len());
fn decode_key_indices(flat_index: usize, key_counts: &[usize]) -> Vec<usize> {
let mut indices = Vec::with_capacity(key_counts.len());
let mut remaining = flat_index; let mut remaining = flat_index;
// Process dimensions in reverse (last dimension varies fastest) for &size in key_sizes.iter().rev() {
for count in key_counts.iter().rev() { indices.push(remaining % size);
indices.push((remaining % count) as usize); remaining /= size;
remaining /= count;
} }
indices.reverse(); // Restore original order indices.reverse();
indices indices
} }
/// Parses a single row from JSON-stat2 format into a DataRow. /// Parses a single row from a JSON-stat2 response into a `DataRow`.
/// ///
/// # Arguments /// Dimension order in the response `id` array determines positional mapping.
/// * `row_index` - Index into the dataset's value array /// The expected order (from API metadata) is:
/// * `dataset` - The complete dataset response /// measure, species, gear, zone, processing, preservation, shipsize, month
/// * `lookup_maps` - Code → label mappings for each dimension
/// ///
/// # Errors /// However, JSON-stat2 `dimension.id` defines the actual order — we read it
/// Returns an error if the row_index exceeds bounds or required lookups are missing. /// dynamically and map by dimension code, not by position assumption.
pub fn parse_row( pub fn parse_row(
row_index: usize, row_index: usize,
dataset: &DataResponse, dataset: &DataResponse,
lookup_maps: &LookupMap, lookup_maps: &LookupMap,
) -> Result<DataRow> { ) -> Result<DataRow> {
let dim_info = &dataset.dataset.dimension; let dim_info = &dataset.dataset.dimension;
let categories = &dim_info.category; let dim_order = &dim_info.id;
let keys = &dim_info.keys;
let value = dataset.dataset.value.get(row_index).copied().flatten(); if dim_order.len() != 8 {
return Err(IngestError::MissingDimension(format!(
"Expected 8 dimensions, got {}. Dimensions: {:?}",
dim_order.len(),
dim_order
)));
}
// Validate bounds
if row_index >= dataset.dataset.value.len() { if row_index >= dataset.dataset.value.len() {
return Err(IngestError::InvalidValueCode(format!( return Err(IngestError::InvalidValueCode(format!(
"Row index {} out of bounds (max {})", "Row index {} out of bounds (max {})",
@@ -181,82 +263,108 @@ pub fn parse_row(
))); )));
} }
// Get category label lists for each dimension (in key order) // Build ordered category lists: for each dimension, extract (code, label)
let category_lists: Vec<Vec<&str>> = keys // pairs sorted by their JSON-stat2 index position.
let category_lists: Vec<Vec<(String, String)>> = dim_order
.iter() .iter()
.filter_map(|k| { .map(|dim_code| {
categories.label.get(k).map(|labels| { let dim = dim_info.dimensions.get(dim_code).ok_or_else(|| {
labels.iter().map(|s| s.as_str()).collect::<Vec<_>>() IngestError::MissingDimension(format!(
}) "Dimension '{}' not found in response",
dim_code
))
})?;
let mut entries: Vec<(String, usize)> = dim
.category
.index
.iter()
.map(|(code, &pos)| (code.clone(), pos))
.collect();
entries.sort_by_key(|(_, pos)| *pos);
let list: Vec<(String, String)> = entries
.into_iter()
.map(|(code, _)| {
let label = dim.category.label.get(&code).cloned().unwrap_or_else(|| {
lookup_maps
.get(dim_code)
.and_then(|m| m.get(&code))
.cloned()
.unwrap_or_else(|| code.clone())
});
(code, label)
}) })
.collect(); .collect();
if category_lists.len() != keys.len() { Ok(list)
return Err(IngestError::MissingDimension( })
format!( .collect::<Result<Vec<_>>>()?;
"Category count mismatch: {} keys vs {} category lists",
keys.len(),
category_lists.len()
)
));
}
// Decode flat index into per-dimension indices let key_sizes: Vec<usize> = category_lists.iter().map(|c| c.len()).collect();
let key_counts: Vec<usize> = category_lists.iter().map(|c| c.len()).collect(); let indices = decode_key_indices(row_index, &key_sizes);
let indices = decode_key_indices(row_index, &key_counts);
// Map indices back to actual dimension values let dimension_values: Vec<(String, String)> = indices
let dimension_values: Vec<&str> = indices
.into_iter() .into_iter()
.zip(category_lists.iter()) .zip(category_lists.iter())
.map(|(idx, list)| list[idx]) .map(|(idx, list)| list[idx].clone())
.collect(); .collect();
// Expected 8 dimensions: [month, species, gear, zone, processing, preservation, shipsize, measure] // Build a lookup from dimension code → (code, label) for this row
const EXPECTED_DIMS: usize = 8; let mut dim_map: HashMap<&str, (String, String)> = HashMap::with_capacity(8);
if dimension_values.len() != EXPECTED_DIMS { for (dim_code, values) in dim_order.iter().zip(dimension_values.iter()) {
return Err(IngestError::MissingDimension(format!( dim_map.insert(dim_code.as_str(), values.clone());
"Expected {} dimensions, got {}",
EXPECTED_DIMS,
dimension_values.len()
)));
} }
let [month_code, species_code, gear_code, zone_code, processing_code, preservation_code, shipsize_code, measure_code]: [&str; 8] = let get = |code: &str| -> (String, String) {
dimension_values.try_into().unwrap(); dim_map
.get(code)
// Helper to safely get label from lookup map .cloned()
fn get_label<'a>(maps: &'a LookupMap, var_name: &str, code: &str) -> &'a str { .unwrap_or_else(|| ("".to_string(), "".to_string()))
maps.get(var_name)
.and_then(|m| m.get(code))
.map(|s| s.as_str())
.unwrap_or(code)
}
let data_row = DataRow {
month: month_code.to_string(),
species_code: species_code.to_string(),
species_label: get_label(lookup_maps, "Art", species_code).to_string(),
gear_code: gear_code.to_string(),
gear_label: get_label(lookup_maps, "Redskab", gear_code).to_string(),
zone_code: zone_code.to_string(),
zone_label: get_label(lookup_maps, "Økonomisk zone", zone_code).to_string(),
processing_code: processing_code.to_string(),
processing_label: get_label(lookup_maps, "Tilstand", processing_code).to_string(),
preservation_code: preservation_code.to_string(),
preservation_label: get_label(lookup_maps, "Konservering", preservation_code).to_string(),
shipsize_code: shipsize_code.to_string(),
shipsize_label: get_label(lookup_maps, "Skibsstørrelse", shipsize_code).to_string(),
measure_code: measure_code.to_string(),
measure_label: get_label(lookup_maps, "Måleenhed", measure_code).to_string(),
value, // Already handled None case above
}; };
Ok(data_row) let (month_code, month_label) = get(DIM_MONTH);
let (species_code, species_label) = get(DIM_SPECIES);
let (gear_code, gear_label) = get(DIM_GEAR);
let (zone_code, zone_label) = get(DIM_ZONE);
let (processing_code, processing_label) = get(DIM_PROCESSING);
let (preservation_code, preservation_label) = get(DIM_PRESERVATION);
let (shipsize_code, shipsize_label) = get(DIM_SHIPSIZE);
let (measure_code, measure_label) = get(DIM_MEASURE);
// Coerce sentinel values to None
let raw_value = dataset.dataset.value[row_index];
let value = match raw_value {
Some(v) if is_sentinel(v) => None,
other => other,
};
debug!(
"Row {}: month={} species={} ({}) measure={} value={:?}",
row_index, month_code, species_code, species_label, measure_code, value
);
Ok(DataRow {
month: month_code,
species_code,
species_label,
gear_code,
gear_label,
zone_code,
zone_label,
processing_code,
processing_label,
preservation_code,
preservation_label,
shipsize_code,
shipsize_label,
measure_code,
measure_label,
value,
})
} }
/// Converts DataRow to Landing struct for database insertion. /// Converts a `DataRow` to a `Landing` struct for database insertion.
/// Removes redundant fields that aren't needed in the fact table.
pub fn data_row_to_landing(row: &DataRow) -> Landing { pub fn data_row_to_landing(row: &DataRow) -> Landing {
Landing { Landing {
month: row.month.clone(), month: row.month.clone(),
@@ -272,139 +380,180 @@ pub fn data_row_to_landing(row: &DataRow) -> Landing {
} }
} }
/// Collects all months available in the metadata response.
/// Used for building incremental ingestion queries.
pub fn extract_available_months(meta: &MetadataResponse) -> Vec<String> {
meta.values
.get("Tid")
.map(|codes| codes.iter().map(|c| c.code.clone()).collect())
.unwrap_or_default()
}
#[cfg(test)] #[cfg(test)]
mod tests { mod tests {
use super::*; use super::*;
use std::sync::OnceLock;
static CLIENT: OnceLock<Client> = OnceLock::new(); fn mock_dataset_response() -> DataResponse {
let mut dimensions = HashMap::new();
fn client() -> &'static Client { dimensions.insert(
CLIENT.get_or_init(Client::new) DIM_MONTH.to_string(),
Dimension {
label: "Mánaður".to_string(),
category: CategoryInfo {
index: HashMap::from([("2015M01".to_string(), 0), ("2015M02".to_string(), 1)]),
label: HashMap::from([
("2015M01".to_string(), "Januar 2015".to_string()),
("2015M02".to_string(), "Februar 2015".to_string()),
]),
},
},
);
dimensions.insert(
DIM_SPECIES.to_string(),
Dimension {
label: "Fiskaslag".to_string(),
category: CategoryInfo {
index: HashMap::from([
("148XXXXXXX00000".to_string(), 0),
("183XXXXXXX00000".to_string(), 1),
]),
label: HashMap::from([
("148XXXXXXX00000".to_string(), "Sild".to_string()),
("183XXXXXXX00000".to_string(), "Þorskur".to_string()),
]),
},
},
);
for dim_code in [
DIM_GEAR,
DIM_ZONE,
DIM_PROCESSING,
DIM_PRESERVATION,
DIM_SHIPSIZE,
] {
dimensions.insert(
dim_code.to_string(),
Dimension {
label: dim_code.to_string(),
category: CategoryInfo {
index: HashMap::from([("TOTAL".to_string(), 0)]),
label: HashMap::from([("TOTAL".to_string(), "Total".to_string())]),
},
},
);
} }
/// Test fixture: mock JSON-stat2 response with Faroese Unicode characters. dimensions.insert(
/// Simulates a response with 2 months × 2 species × 2 measures = 8 values. DIM_MEASURE.to_string(),
fn mock_dataset_response() -> DataResponse { Dimension {
label: "Mát".to_string(),
category: CategoryInfo {
index: HashMap::from([("MASS".to_string(), 0), ("VALUE".to_string(), 1)]),
label: HashMap::from([
("MASS".to_string(), "Nøgd".to_string()),
("VALUE".to_string(), "Virði".to_string()),
]),
},
},
);
// Dimension order: measure, species, gear, zone, processing,
// preservation, shipsize, month
// Sizes: 2, 2, 1, 1, 1, 1, 1, 2 = 8 values
DataResponse { DataResponse {
dataset: Dataset { dataset: Dataset {
// Keys specify cardinality per dimension (2 months, 2 species, ..., 2 measures) dimension: DimInfo {
keys: vec![ id: vec![
"2".to_string(), // Time (2 months) DIM_MEASURE.to_string(),
"2".to_string(), // Art (2 species) DIM_SPECIES.to_string(),
"1".to_string(), // Redskab (1 gear - TOTAL) DIM_GEAR.to_string(),
"1".to_string(), // Zone (1 zone - TOTAL) DIM_ZONE.to_string(),
"1".to_string(), // Tilstand (1 - TOTAL) DIM_PROCESSING.to_string(),
"1".to_string(), // Konservering (1 - TOTAL) DIM_PRESERVATION.to_string(),
"1".to_string(), // Skibsstørrelse (1 - TOTAL) DIM_SHIPSIZE.to_string(),
"2".to_string(), // Måleenhed (2 - MASS, VALUE) DIM_MONTH.to_string(),
], ],
category: CategoryInfo { size: vec![2, 2, 1, 1, 1, 1, 1, 2],
label: HashMap::from([ dimensions,
(
"0".to_string(),
vec!["2015M01".to_string(), "2015M02".to_string()],
),
(
"1".to_string(),
vec!["Sild".to_string(), "Þorskur".to_string()],
),
("2".to_string(), vec!["TOTAL".to_string()]),
("3".to_string(), vec!["TOTAL".to_string()]),
("4".to_string(), vec!["TOTAL".to_string()]),
("5".to_string(), vec!["TOTAL".to_string()]),
("6".to_string(), vec!["TOTAL".to_string()]),
(
"7".to_string(),
vec!["MASS".to_string(), "VALUE".to_string()],
),
]),
index: None,
}, },
value: vec![ value: vec![
Some(1234.5), // 2015M01, Sild, TOTAL..., MASS Some(1234.5), // MASS, Sild, ..., 2015M01
Some(2345.6), // 2015M01, Sild, TOTAL..., VALUE Some(2345.6), // MASS, Sild, ..., 2015M02
Some(-1.0), // 2015M01, Þorskur, TOTAL..., MASS (marker) Some(-1.0), // MASS, Þorskur, ..., 2015M01 (sentinel)
Some(3456.7), // 2015M01, Þorskur, TOTAL..., VALUE Some(3456.7), // MASS, Þorskur, ..., 2015M02
Some(4567.8), // 2015M02, Sild, TOTAL..., MASS Some(4567.8), // VALUE, Sild, ..., 2015M01
Some(5678.9), // 2015M02, Sild, TOTAL..., VALUE Some(5678.9), // VALUE, Sild, ..., 2015M02
None, // 2015M02, Þorskur, TOTAL..., MASS (missing) None, // VALUE, Þorskur, ..., 2015M01 (null)
Some(6789.0), // 2015M02, Þorskur, TOTAL..., VALUE Some(6789.0), // VALUE, Þorskur, ..., 2015M02
], ],
status: vec![],
}, },
} }
} }
/// Test fixture: mock lookup maps with Faroese labels.
fn mock_lookup_maps() -> LookupMap { fn mock_lookup_maps() -> LookupMap {
let mut maps = LookupMap::new(); let mut maps = LookupMap::new();
maps.insert( maps.insert(
"0".to_string(), DIM_MONTH.to_string(),
HashMap::from([ HashMap::from([
("2015M01".to_string(), "Januar 2015".to_string()), ("2015M01".to_string(), "Januar 2015".to_string()),
("2015M02".to_string(), "Februar 2015".to_string()), ("2015M02".to_string(), "Februar 2015".to_string()),
]), ]),
); );
maps.insert( maps.insert(
"1".to_string(), DIM_SPECIES.to_string(),
HashMap::from([ HashMap::from([
("Sild".to_string(), "Sild".to_string()), ("148XXXXXXX00000".to_string(), "Sild".to_string()),
("Þorskur".to_string(), "Þorskur".to_string()), ("183XXXXXXX00000".to_string(), "Þorskur".to_string()),
]), ]),
); );
maps.insert( maps.insert(
"7".to_string(), DIM_MEASURE.to_string(),
HashMap::from([ HashMap::from([
("MASS".to_string(), "Kilo".to_string()), ("MASS".to_string(), "Nøgd".to_string()),
("VALUE".to_string(), "Krónur".to_string()), ("VALUE".to_string(), "Virði".to_string()),
]), ]),
); );
maps maps
} }
#[tokio::test] #[test]
async fn test_build_query_all_wildcards() { fn test_build_query_uses_correct_dimension_codes() {
let months = vec!["2015M01".to_string(), "2015M02".to_string()]; let months = vec!["2015M01".to_string(), "2015M02".to_string()];
let query = build_query(&months); let query = build_query(&months);
assert_eq!(query.language, PX_WEB_LANGUAGE); assert_eq!(query.response.format, "json-stat2");
assert_eq!(query.query.len(), 8); // All 8 dimensions assert_eq!(query.query.len(), 8);
// Check that species is wildcard assert_eq!(query.query[0].code, DIM_MONTH);
let species_query = query.query.iter().find(|q| q.id == "Art").unwrap(); assert_eq!(query.query[0].selection.filter, "item");
assert_eq!(species_query.values, vec!["*".to_string()]); assert_eq!(query.query[0].selection.values.len(), 2);
// Check that measures includes both MASS and VALUE assert_eq!(query.query[1].code, DIM_SPECIES);
let measure_query = query.query.iter().find(|q| q.id == "Måleenhed").unwrap(); assert_eq!(query.query[1].selection.filter, "all");
assert!(measure_query.values.contains(&"MASS".to_string()));
assert!(measure_query.values.contains(&"VALUE".to_string())); assert_eq!(query.query[7].code, DIM_MEASURE);
assert!(
query.query[7]
.selection
.values
.contains(&"MASS".to_string())
);
assert!(
query.query[7]
.selection
.values
.contains(&"VALUE".to_string())
);
} }
#[test] #[test]
fn test_decode_key_indices_basic() { fn test_decode_key_indices_basic() {
// 2 × 2 × 2 = 8 combinations let key_sizes = vec![2, 2, 2];
let key_counts = vec![2, 2, 2]; assert_eq!(decode_key_indices(0, &key_sizes), vec![0, 0, 0]);
assert_eq!(decode_key_indices(7, &key_sizes), vec![1, 1, 1]);
assert_eq!(decode_key_indices(3, &key_sizes), vec![0, 1, 1]);
}
// Index 0 should give [0, 0, 0] #[test]
let indices = decode_key_indices(0, &key_counts); fn test_decode_key_indices_single_dim() {
assert_eq!(indices, vec![0, 0, 0]); let key_sizes = vec![5];
assert_eq!(decode_key_indices(0, &key_sizes), vec![0]);
// Index 7 should give [1, 1, 1] assert_eq!(decode_key_indices(4, &key_sizes), vec![4]);
let indices = decode_key_indices(7, &key_counts);
assert_eq!(indices, vec![1, 1, 1]);
// Index 3 should give [0, 1, 1] (middle combination)
let indices = decode_key_indices(3, &key_counts);
assert_eq!(indices, vec![0, 1, 1]);
} }
#[test] #[test]
@@ -412,43 +561,52 @@ mod tests {
let dataset = mock_dataset_response(); let dataset = mock_dataset_response();
let lookup_maps = mock_lookup_maps(); let lookup_maps = mock_lookup_maps();
// Parse first row (index 0) let row = parse_row(0, &dataset, &lookup_maps).expect("parse failed");
let row = parse_row(0, &dataset, &lookup_maps).expect("Failed to parse row");
// Verify Faroese characters survive intact
assert_eq!(row.month, "2015M01"); assert_eq!(row.month, "2015M01");
assert_eq!(row.species_code, "Sild"); assert_eq!(row.species_code, "148XXXXXXX00000");
assert!(row.species_label.contains("Sild")); assert_eq!(row.species_label, "Sild");
assert_eq!(row.measure_code, "MASS"); assert_eq!(row.measure_code, "MASS");
assert_eq!(row.measure_label, "Nøgd");
// Verify value is preserved assert_eq!(row.value, Some(1234.5));
assert!(row.value.is_some());
assert_eq!(row.value.unwrap(), 1234.5);
} }
#[test] #[test]
fn test_parse_row_second_species() { fn test_parse_row_thorn_character() {
let dataset = mock_dataset_response(); let dataset = mock_dataset_response();
let lookup_maps = mock_lookup_maps(); let lookup_maps = mock_lookup_maps();
// Row index 2 should be Þorskur (second species, first time, MASS) // Row 2: MASS, Þorskur, ..., 2015M01
let row = parse_row(2, &dataset, &lookup_maps).expect("Failed to parse row"); let row = parse_row(2, &dataset, &lookup_maps).expect("parse failed");
// Verify the special character survives assert_eq!(row.species_code, "183XXXXXXX00000");
assert_eq!(row.species_code, "Þorskur"); assert_eq!(row.species_label, "Þorskur");
assert!(row.species_label.contains("Þ")); assert!(row.species_label.contains('Þ'));
assert!(row.species_label.contains("orskur"));
} }
#[test] #[test]
fn test_parse_row_missing_value_handling() { fn test_parse_row_sentinel_coerced_to_none() {
let dataset = mock_dataset_response(); let dataset = mock_dataset_response();
let lookup_maps = mock_lookup_maps(); let lookup_maps = mock_lookup_maps();
// Row index 6 has None value (missing data for 2015M02, Þorskur, MASS) // Row 2 has -1.0 sentinel
let row = parse_row(6, &dataset, &lookup_maps).expect("Failed to parse row"); let row = parse_row(2, &dataset, &lookup_maps).expect("parse failed");
assert!(
row.value.is_none(),
"Sentinel -1.0 must be None, got {:?}",
row.value
);
}
#[test]
fn test_parse_row_explicit_null() {
let dataset = mock_dataset_response();
let lookup_maps = mock_lookup_maps();
// Row 6 has None
let row = parse_row(6, &dataset, &lookup_maps).expect("parse failed");
// Missing values should become None, not panic
assert!(row.value.is_none()); assert!(row.value.is_none());
} }
@@ -457,48 +615,81 @@ mod tests {
let dataset = mock_dataset_response(); let dataset = mock_dataset_response();
let lookup_maps = mock_lookup_maps(); let lookup_maps = mock_lookup_maps();
// Requesting out-of-bounds index should return error
let result = parse_row(100, &dataset, &lookup_maps); let result = parse_row(100, &dataset, &lookup_maps);
assert!(result.is_err()); assert!(result.is_err());
if let Err(IngestError::InvalidValueCode(msg)) = result { match result {
assert!(msg.contains("out of bounds")); Err(IngestError::InvalidValueCode(msg)) => assert!(msg.contains("out of bounds")),
} else { _ => panic!("Expected InvalidValueCode"),
panic!("Expected InvalidValueCode error");
} }
} }
#[test]
fn test_parse_row_wrong_dimension_count() {
let mut dataset = mock_dataset_response();
dataset.dataset.dimension.id.pop();
dataset.dataset.dimension.size.pop();
let lookup_maps = mock_lookup_maps();
let result = parse_row(0, &dataset, &lookup_maps);
assert!(result.is_err());
}
#[test] #[test]
fn test_data_row_to_landing() { fn test_data_row_to_landing() {
let dataset = mock_dataset_response(); let dataset = mock_dataset_response();
let lookup_maps = mock_lookup_maps(); let lookup_maps = mock_lookup_maps();
let row = parse_row(0, &dataset, &lookup_maps).expect("Failed to parse row"); let row = parse_row(0, &dataset, &lookup_maps).expect("parse failed");
let landing = data_row_to_landing(&row); let landing = data_row_to_landing(&row);
// Verify core fields are copied
assert_eq!(landing.month, row.month); assert_eq!(landing.month, row.month);
assert_eq!(landing.species_code, row.species_code); assert_eq!(landing.species_code, row.species_code);
assert_eq!(landing.species_label, row.species_label);
assert_eq!(landing.measure_code, row.measure_code); assert_eq!(landing.measure_code, row.measure_code);
assert_eq!(landing.value, row.value); assert_eq!(landing.value, row.value);
}
// Verify lookup fields are preserved #[test]
assert_eq!(landing.species_label, row.species_label); fn test_faroese_unicode_all_rows() {
assert_eq!(landing.gear_code, row.gear_code); let dataset = mock_dataset_response();
assert_eq!(landing.zone_code, row.zone_code); let lookup_maps = mock_lookup_maps();
for i in 0..8 {
let row = parse_row(i, &dataset, &lookup_maps);
assert!(row.is_ok(), "Row {} failed", i);
if let Ok(r) = row {
assert!(
!r.species_label.contains('\u{FFFD}'),
"Replacement char in row {}",
i
);
}
}
}
#[test]
fn test_is_sentinel() {
assert!(is_sentinel(-1.0));
assert!(is_sentinel(f64::NAN));
assert!(!is_sentinel(0.0));
assert!(!is_sentinel(1234.5));
assert!(!is_sentinel(0.001));
} }
#[test] #[test]
fn test_extract_available_months() { fn test_extract_available_months() {
let meta = MetadataResponse { let meta = MetadataResponse {
variables: vec![], title: Some("test".to_string()),
values: HashMap::from([( variables: vec![VariableMeta {
"Tid".to_string(), code: DIM_MONTH.to_string(),
vec![ text: "Mánaður".to_string(),
ValueMeta { code: "2015M01".to_string(), text: "Jan 2015".to_string() }, role: None,
ValueMeta { code: "2015M02".to_string(), text: "Feb 2015".to_string() }, values: vec!["2015M01".to_string(), "2015M02".to_string()],
], value_texts: vec!["Jan 2015".to_string(), "Feb 2015".to_string()],
)]), extra: HashMap::new(),
}],
}; };
let months = extract_available_months(&meta); let months = extract_available_months(&meta);
@@ -506,4 +697,45 @@ mod tests {
assert!(months.contains(&"2015M01".to_string())); assert!(months.contains(&"2015M01".to_string()));
assert!(months.contains(&"2015M02".to_string())); assert!(months.contains(&"2015M02".to_string()));
} }
#[test]
fn test_extract_available_months_no_month_var() {
let meta = MetadataResponse {
title: None,
variables: vec![VariableMeta {
code: DIM_SPECIES.to_string(),
text: "Fiskaslag".to_string(),
role: None,
values: vec!["TOTAL".to_string()],
value_texts: vec!["Tils.".to_string()],
extra: HashMap::new(),
}],
};
let months = extract_available_months(&meta);
assert!(months.is_empty());
}
#[test]
#[should_panic(expected = "requires at least one month")]
fn test_build_query_empty_panics() {
let _ = build_query(&[]);
}
#[test]
fn test_variable_meta_lookup_map() {
let var = VariableMeta {
code: "measure".to_string(),
text: "mát".to_string(),
role: None,
values: vec!["MASS".to_string(), "VALUE".to_string()],
value_texts: vec!["Nøgd".to_string(), "Virði".to_string()],
extra: HashMap::new(),
};
let map = var.lookup_map();
assert_eq!(map.get("MASS"), Some(&"Nøgd".to_string()));
assert_eq!(map.get("VALUE"), Some(&"Virði".to_string()));
assert_eq!(map.len(), 2);
}
} }
+1 -1
View File
@@ -1,3 +1,3 @@
fn main() { fn main() {
println!("Hello, world!"); println!("hagfish — not yet wired. Run tests with: cargo test");
} }
+140 -36
View File
@@ -20,63 +20,132 @@ impl Default for Config {
Self { Self {
duckdb_path: "hagfish.db".to_string(), duckdb_path: "hagfish.db".to_string(),
bind_address: "127.0.0.1:8090".to_string(), bind_address: "127.0.0.1:8090".to_string(),
data_source_url: "https://statbank.hagstova.fo/data/api/table/fisknv_md".to_string(), data_source_url: "https://statbank.hagstova.fo/api/v1/fo/H2/VV/VV01/fisknv_md.px"
.to_string(),
log_file_path: Some("hagfish.log".to_string()), log_file_path: Some("hagfish.log".to_string()),
} }
} }
} }
// ---------------------------------------------------------------------------
// Metadata types (GET response)
// ---------------------------------------------------------------------------
/// Metadata response from PX-Web API GET request. /// Metadata response from PX-Web API GET request.
/// Contains variable definitions and value codes with labels. ///
/// Confirmed structure from live API:
/// ```json
/// {
/// "title": "AVR01010 ...",
/// "variables": [
/// {
/// "code": "measure",
/// "text": "mát",
/// "role": null,
/// "values": ["MASS", "VALUE"],
/// "valueTexts": ["Nøgd", "Virði"]
/// }
/// ]
/// }
/// ```
#[derive(Debug, Clone, Deserialize, Serialize)] #[derive(Debug, Clone, Deserialize, Serialize)]
pub struct MetadataResponse { pub struct MetadataResponse {
#[serde(default)]
pub title: Option<String>,
pub variables: Vec<VariableMeta>, pub variables: Vec<VariableMeta>,
pub values: HashMap<String, Vec<ValueMeta>>,
} }
/// Variable metadata describing a dimension in the PX-Web table. /// Variable metadata describing a dimension in the PX-Web table.
///
/// PxWeb v1 uses parallel string arrays for codes and labels.
/// The `role` field is null on this endpoint — do not rely on it.
#[derive(Debug, Clone, Deserialize, Serialize)] #[derive(Debug, Clone, Deserialize, Serialize)]
pub struct VariableMeta { pub struct VariableMeta {
pub id: String, /// Variable identifier used in queries (e.g. "month", "measure",
/// "Species (ASFIS2022)")
pub code: String,
/// Human-readable label in the queried language
#[serde(default)]
pub text: String, pub text: String,
pub role: String, /// Role classification — null on this endpoint
/// Position in the key array (0-indexed) #[serde(default)]
#[serde(rename = "keyPosition", skip_serializing_if = "Option::is_none")] pub role: Option<String>,
pub key_position: Option<usize>, /// Code values for this variable (parallel with valueTexts)
#[serde(default)]
pub values: Vec<String>,
/// Display labels for each value (parallel with values)
#[serde(default, rename = "valueTexts")]
pub value_texts: Vec<String>,
/// Catch-all for unknown fields
#[serde(flatten)]
pub extra: HashMap<String, serde_json::Value>,
} }
/// Value metadata: code → display label mapping for a variable. impl VariableMeta {
#[derive(Debug, Clone, Deserialize, Serialize)] /// Build a code→label lookup map from the parallel arrays.
pub struct ValueMeta { pub fn lookup_map(&self) -> HashMap<String, String> {
pub code: String, self.values
pub text: String, .iter()
.zip(self.value_texts.iter())
.map(|(code, label)| (code.clone(), label.clone()))
.collect()
}
} }
// ---------------------------------------------------------------------------
// Query types (POST request body)
// ---------------------------------------------------------------------------
/// Query body sent to PX-Web API POST endpoint. /// Query body sent to PX-Web API POST endpoint.
///
/// Confirmed format:
/// ```json
/// {
/// "query": [{"code": "month", "selection": {"filter": "item", "values": [...]}}],
/// "response": {"format": "json-stat2"}
/// }
/// ```
#[derive(Debug, Clone, Serialize)] #[derive(Debug, Clone, Serialize)]
pub struct Query { pub struct Query {
pub query: Vec<QueryItem>, pub query: Vec<QueryItem>,
pub language: String, pub response: QueryResponse,
} }
/// Single query item representing a dimension selection. /// Single query item representing a dimension selection.
#[derive(Debug, Clone, Serialize)] #[derive(Debug, Clone, Serialize)]
pub struct QueryItem { pub struct QueryItem {
pub id: String, pub code: String,
pub selection: Selection,
}
/// Selection within a query item.
#[derive(Debug, Clone, Serialize)]
pub struct Selection {
/// "item" for explicit values, "all" for wildcard, "top" for top-N
pub filter: String,
/// Values to select. For "all" filter, use ["*"].
pub values: Vec<String>, pub values: Vec<String>,
} }
/// Selection helper for building queries programmatically. /// Response format specification.
#[derive(Debug, Clone)] #[derive(Debug, Clone, Serialize)]
pub struct Selection { pub struct QueryResponse {
/// Dimension name (matches variable.id from metadata) pub format: String,
pub dimension: String,
/// Codes to include, or "*" for all
pub codes: Vec<String>,
} }
impl Default for QueryResponse {
fn default() -> Self {
Self {
format: "json-stat2".to_string(),
}
}
}
// ---------------------------------------------------------------------------
// Data response types (JSON-stat2)
// ---------------------------------------------------------------------------
/// Raw data response from PX-Web API POST request. /// Raw data response from PX-Web API POST request.
/// JSON-stat2 format with metadata and data sections.
#[derive(Debug, Clone, Deserialize, Serialize)] #[derive(Debug, Clone, Deserialize, Serialize)]
pub struct DataResponse { pub struct DataResponse {
pub dataset: Dataset, pub dataset: Dataset,
@@ -87,24 +156,52 @@ pub struct DataResponse {
pub struct Dataset { pub struct Dataset {
pub dimension: DimInfo, pub dimension: DimInfo,
pub value: Vec<Option<f64>>, pub value: Vec<Option<f64>>,
/// Status codes per value (optional)
#[serde(default, skip_serializing_if = "Vec::is_empty")]
pub status: Vec<String>,
} }
/// Dimension metadata describing key layout. /// Dimension metadata in JSON-stat2 format.
///
/// Uses a flat `id` array for ordering and a `size` array for cardinality.
/// Each dimension is keyed by its code in the `dimensions` map.
#[derive(Debug, Clone, Deserialize, Serialize)] #[derive(Debug, Clone, Deserialize, Serialize)]
pub struct DimInfo { pub struct DimInfo {
#[serde(rename = "key")] /// Ordered dimension IDs (defines cube layout)
pub keys: Vec<String>, #[serde(default)]
pub id: Vec<String>,
/// Size of each dimension
#[serde(default)]
pub size: Vec<usize>,
/// One entry per dimension, keyed by dimension code
#[serde(flatten)]
pub dimensions: HashMap<String, Dimension>,
}
/// A single dimension in the JSON-stat2 response.
#[derive(Debug, Clone, Deserialize, Serialize)]
pub struct Dimension {
/// Display label
#[serde(default)]
pub label: String,
/// Category info with index and label maps
pub category: CategoryInfo, pub category: CategoryInfo,
} }
/// Category info containing value lists for each dimension. /// Category info containing index and label maps.
#[derive(Debug, Clone, Deserialize, Serialize)] #[derive(Debug, Clone, Deserialize, Serialize)]
pub struct CategoryInfo { pub struct CategoryInfo {
pub label: HashMap<String, Vec<String>>, /// Maps category code → numeric position in the dimension
#[serde(skip_serializing_if = "Option::is_none")] pub index: HashMap<String, usize>,
pub index: Option<HashMap<String, Vec<usize>>>, /// Maps category code → display label
#[serde(default)]
pub label: HashMap<String, String>,
} }
// ---------------------------------------------------------------------------
// Internal data models
// ---------------------------------------------------------------------------
/// Decoded row of landing data with labeled dimensions. /// Decoded row of landing data with labeled dimensions.
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
pub struct DataRow { pub struct DataRow {
@@ -126,7 +223,7 @@ pub struct DataRow {
pub value: Option<f64>, pub value: Option<f64>,
} }
/// Structured representation of a single landing record. /// Structured representation of a single landing record for DB insertion.
#[derive(Debug, Clone)] #[derive(Debug, Clone)]
pub struct Landing { pub struct Landing {
pub month: String, pub month: String,
@@ -141,6 +238,10 @@ pub struct Landing {
pub value: Option<f64>, pub value: Option<f64>,
} }
// ---------------------------------------------------------------------------
// Errors
// ---------------------------------------------------------------------------
/// Error types for ingestion module. /// Error types for ingestion module.
#[derive(Debug, thiserror::Error)] #[derive(Debug, thiserror::Error)]
pub enum IngestError { pub enum IngestError {
@@ -153,18 +254,21 @@ pub enum IngestError {
#[error("Missing dimension in response: {0}")] #[error("Missing dimension in response: {0}")]
MissingDimension(String), MissingDimension(String),
#[error("Invalid value code: {0}")] #[error("Invalid value: {0}")]
InvalidValueCode(String), InvalidValueCode(String),
#[error("API returned empty data set")] #[error("API returned empty data set")]
EmptyDataset, EmptyDataset,
#[error("Unicode decode error: {0}")] #[error("Dimension '{0}' not found in metadata")]
UnicodeError(String), DimensionNotFound(String),
} }
/// Lookup map for decoding key arrays into labeled values. // ---------------------------------------------------------------------------
/// Keyed by dimension name, contains code → label mappings. // Type aliases
// ---------------------------------------------------------------------------
/// Lookup map: dimension_code → (value_code → display_label).
pub type LookupMap = HashMap<String, HashMap<String, String>>; pub type LookupMap = HashMap<String, HashMap<String, String>>;
/// Result alias using custom error type. /// Result alias using custom error type.