Snapshot of where the project stands. Companion to RESEARCH.md (scouting) and PLAN.md (design).
End-to-end pipeline from HMIP endpoint → queryable parquet archive → portable analyst-facing data mart:
catalogue (285 entries)
↓
bulk_pull(geographies) → asyncio.Semaphore(5) → fetch_table_async()
↓ ↓
↓ data/raw/{survey}/{table_id}/{geo}.csv
↓ ↓
↓ data/logs/{label}_{ts}.jsonl (one record / attempt)
↓
build_parquet → tidy() → data/clean/{survey}/{table_id}.parquet
↓
DuckDB ad-hoc (reads parquet directly — no warehouse layer)
↓
build_dmt_rental.py → data/marts/cmhc_rental.duckdb (Ontario rental, star + materialized metrics)
Plus a parallel static-data-tables lane (non-HMIP xlsx surface), converging on the same long schema:
build_static_catalogue.py (+--render) → data/static_catalogue.json (128 asset URLs)
↓
download_static.py → data/raw/static/{section}/{file}.xlsx
↓
build_static_parquet.py → matrix engine (MatrixSpec + catalogue provenance, via cmhc.wide)
↓ ↓
↓ data/clean/static/{table_id}.parquet (source='static')
↓
(future) source-agnostic mart builder → unions data/clean/ (HMIP) + data/clean/static/
| File | Purpose |
|---|---|
catalogue.py |
285 (survey, series, dimension, breakdown, geo_filter) → table_id entries. Ported from mountainMath/cmhc's R package. |
geographies.py |
Canada + 13 provinces/territories + 153 CMAs + 574 Ontario CSDs (168 CMA-member subset) + 2,382 Ontario CTs. Geography.geography_id is a string (preserves leading zeros in METCODE and the dotted compound CT id). Also exports normalize_name() for slash/hyphen drift across StatCan ↔ CMHC name spaces. |
hmip.py |
Sync fetch_table + async fetch_table_async (shared _build_form, separate httpx.Client / AsyncClient). Exponential-backoff retries on 5xx + transport errors. is_empty_response() detects HMIP "No data available" / "archived" sentinels. |
tidy.py |
HMIP wide CSV → long polars DataFrame. Handles reliability codes, suppression sentinels, snapshot vs time-series shapes. _parse_period converts 'Feb 1990' / '1990 March' / '1991/Q1' to start-of-period date. Snapshot tables get their period from the CSV subtitle via _extract_subtitle_period. Single-geo query rows (empty first cell) preserved as sub_geography=null. Melt delegated to wide.py. |
wide.py |
Pure shape primitive: wide_to_long(df, index, is_reliability, parse_value) melts a wide matrix (value columns each optionally trailed by a reliability column) into long (index, category, value, reliability). Shared by tidy.py (HMIP) and the static engine; the reliability-detection and value-parsing rules are injected, so neither surface special-cases the other. |
validity.py |
is_valid_for_geo(table, geo) filters job lists for Canada/Province/CMA/CSD/CT geos (HMIP silently returns garbage for invalid combos). BROKEN_TABLE_IDS denylist for known-bad table_ids. |
bulk.py |
The async orchestrator. bulk_pull(geographies, *, label, surveys=None, concurrency=None, refresh_empty_days=None, exclude_tables=None) walks catalogue × geos, filters via is_valid_for_geo (and optional survey allowlist / exclude_tables run-time skip), fetches in parallel under a semaphore, writes CSVs / empty markers, emits per-attempt JSONL log. exclude_tables is a this-run-only skip (e.g. tables HMIP 500s on at a given geo level) — it does not touch the catalogue or validity. |
config.py |
PROJECT_ROOT, RAW_DIR, CLEAN_DIR, EMPTY_DIR, LOG_DIR, STATIC_RAW_DIR, STATIC_CATALOGUE, REQUEST_DELAY, CONCURRENCY. |
static/ |
Non-HMIP static-table harvest. schema.py — the shared long-format contract (COLUMNS = HMIP schema + source). catalogue.py — accessor over static_catalogue.json (slug→table_id, section→survey, asset meta, page title). matrix.py — the configurable engine: MatrixSpec/Sheets + run() (header detection, sheet-name-as-dimension, period parsing, sentinels; melt via wide.py). specs.py — typed recipe registry, slug→MatrixSpec. runner.py — parse(table_id, path). |
| File | Purpose |
|---|---|
pull_canada_and_provinces.py |
Pull at Canada + provincial scope. |
pull_cmas.py |
--province NAME (filters by cma_uid prefix), --surveys, --concurrency, --refresh-empty-days. |
pull_csds.py |
Ontario CSDs. Defaults to ~168 CMA-member subset; --all for all ~574. Same --surveys / --concurrency / --refresh-empty-days flags. |
pull_cts.py |
Ontario CTs (~2,382). Same flags as pull_csds.py (no --all; CTs only exist inside CMAs), plus --exclude-tables IDS to skip tables HMIP 500s on at CT level (run-time only — see DATA_DISCOVERY.md 2026-06-17). |
extract_geo_lookups.py |
One-shot: download .rda lookup tables from mountainMath/cmhc, write Ontario-filtered CSVs into src/cmhc/data/. Re-run when the R package updates. |
build_boundaries.py |
One-shot: download Statistics Canada 2021 cartographic boundary files (CSD + CT), filter to Ontario, reproject to WGS84, topology-simplify via topojson package, write GeoJSON to data/clean/boundaries_*.geojson. |
build_parquet.py |
Walk raw, tidy, concat by table_id, write parquet. Mtime-idempotent — full rebuild requires rm -rf data/clean/ (needed when tidy.py schema changes). |
build_dmt_rental.py |
Tidy parquet → single-file DuckDB data mart for Ontario rental (Rms + Srms), Province→CMA→CSD→CT. Star schema + materialized metric tables. ~38 MB output. See DATAMART.md. |
example_queries.py |
DuckDB query demos against the cleaned parquet. |
build_static_catalogue.py |
Discover static-data-table .xlsx/.xls assets on cmhc-schl.gc.ca. --render (Playwright, scrape group) captures JS-injected downloads. 128 of 136 pages have a captured asset. |
download_static.py |
Fetch every catalogued static asset → data/raw/static/{section}/. Idempotent (skips existing unless --force). |
build_static_parquet.py |
Parse each spec'd static table via the matrix engine → data/clean/static/{table_id}.parquet. Mtime-idempotent; only builds tables present in specs.py. |
probe_table.py |
Diagnostic single-table prober. probe_table.py <table_id> --geo <name> tries bare + catalogue + leave-one-out filter variants; identifies stale catalogue filters. See DATA_DISCOVERY.md. |
| File | Source | Rows |
|---|---|---|
cmas.csv |
cmhc_cma_translation_data.rda |
153 (Canada-wide) |
csds_ontario.csv |
cmhc_csd_translation_data.rda filtered to PR=35 |
574 |
csds_ontario_cma_members.csv |
cmhc_csd_translation_data_2023.rda filtered to PR=35, joined to name/type |
168 |
cts_ontario.csv |
cmhc_ct_translation_data.rda filtered to PR=35, with precomputed GeographyId = METCODE + NBHDCODE + CMHC_CT |
2,382 (15 source rows with whitespace-only CTUID dropped) |
GeoJSON polygons from Statistics Canada 2021 Cartographic Boundary Files, simplified topologically (preserves shared edges) via the topojson package. Used for choropleth rendering.
| File | Features | Size | Join key |
|---|---|---|---|
boundaries_csd_ontario.geojson |
577 CSDs (PR=35) | 1.9 MB | CSDUID (568 match csds_ontario.csv) |
boundaries_ct_ontario.geojson |
2,533 CTs (PR=35) | 2.0 MB | CTUID (2,265 match cts_ontario.csv; mismatch is 2016 vs 2021 CT vintage) |
Raw zips cached at data/raw/boundaries/ to avoid re-downloading (~50 MB). Build with uv run python scripts/build_boundaries.py.
63 unit tests covering catalogue, geographies (incl. Ontario CSD + CT lookups), hmip, tidy (period parsing, snapshot CSV shape, subtitle period extraction, single-geo row preservation), validity, and the static matrix engine (the three layout families, reliability columns, divider/footnote dropping, registry + catalogue wiring). All green.
Coverage spans Canada + 13 provinces + 42 Ontario CMAs + 147 Ontario CSDs + 1,589 Ontario CTs. Row counts change every pull — get live numbers from the data, not this doc: SELECT count(*) FROM 'data/clean/**/*.parquet'.
| Survey | Coverage |
|---|---|
| Rms | Vacancy / Availability / Rent / Universe — by Bedroom Type, Year of Construction, Structure Size, Rent Range, Rent Quartile, plus Summary Statistics. Canada + provinces + Ontario CMAs + Ontario CMA-member CSDs + Ontario CTs (15 tract-publishing tables; 2026-06-17). Massive 2026-06-09 recovery: bedroom-filter + tidy fixes plus --refresh-empty-days 0 re-pull at all sub-CMA levels recovered rows hidden by stale empty markers. |
| Scss | Starts, Completions, Intended Market, Unabsorbed Inventory — snapshot, Canada time-series, + Ontario CMAs |
| Census | Census-derived counts (Ontario CMAs) |
| Srms | Secondary Rental Market (condo / suite) — 8 Ontario CMAs that publish Srms: Barrie, Hamilton, Kitchener – Cambridge – Waterloo, London, Ottawa, St. Catharines – Niagara, Toronto, Windsor. The other 35 Ontario CMAs return empty (confirmed 2026-06-09). |
| Core Housing Need | Core housing need indicators (Ontario CMAs) |
| Seniors | Seniors housing — Ontario CMAs (incl. snapshot-shape tables) |
data/raw/ holds the raw CSVs plus empty markers — (table, geo) combos HMIP confirmed have no data, so we don't re-fetch them. The 2026-06-17 Ontario CT pull (15 Rms tables × 2,382 CTs) is heavily confidentiality-suppressed, so most CT cells land as empty markers.
All 128 catalogued assets downloaded to data/raw/static/ (~50 MB; 118 xlsx + 10 xls). 18 tables spec'd and built to data/clean/static/ (source='static'):
| Section | Spec'd / on disk | Notes |
|---|---|---|
| Mortgage and Debt | 1 / 18 | Mortgage delinquency rate (Equifax; Canada/prov/CMA, 2012Q3–2025Q4). Rest of the section queued. |
| Household Characteristics | 17 / 52 | Household counts, ownership rates, real median/average income (by tenure), core-housing-need counts/incidence. |
| Rental Market | 0 / 20 | Some overlap HMIP Rms; uniques (seniors survey, non-resident ownership, percentile rents) queued. |
| Housing Market Data | 0 / 38 | Deliberately skipped — duplicates HMIP Scss (starts/completions/absorption). |
Why only 18 of the parseable-looking files are spec'd: a flat-matrix parse raises no error on a multi-dimensional table but produces wrong data, so each candidate is screened by a correctness gate (no all-null category; no duplicate (geography, period, category) keys) before earning a spec. Of 52 household-characteristics files: 17 clean, ~20 multi-dimensional (Tenure / Age group / Quintile breakdown — await a multi-dimension engine mode), 15 structural (no detectable header, or by-geography sheets). See DATA_DISCOVERY.md 2026-06-13.
data/marts/cmhc_rental.duckdb — Ontario rental extract for analyst handoff. Star schema + 25 materialized metric tables, ~39 MB. CT level added 2026-06-17 (Rms only; CT rental is heavily confidentiality-suppressed). Live counts (observations, canonical, suppressed, geographies) and the coverage statement are in the mart's _meta table — query it rather than restating here. See DATAMART.md. Rebuild: uv run python scripts/build_dmt_rental.py.
Correctness pass (2026-06-18, from a hands-on mart evaluation): (#1) the same value published through 2–5 table_id paths silently triplicated the headline cross-section query — added is_canonical (star keeps all paths; materialized tables filter to one row per logical cell). (#3) sort_order was alphabetical (Studio after Total) and absent from the metric tables — replaced with an explicit per-dimension order and joined in. (#4) reliability IS NULL was overloaded (suppressed vs no-info) and silently dropped ~335k valid rows from reliability IN ('a','b') — non-suppressed unrated rows now carry the 'n/a' sentinel, so NULL ⇔ is_suppressed. (#2) Summary Statistics (a repackaging of metrics 1–6) dropped from the mart. Regression tests in tests/test_dmt_rental.py.
Parquet schema (uniform):
period date | null start-of-period date. Populated for time-series (from the
data column) and for snapshots (from the CSV subtitle, e.g.
'October 2025 …'). Null only when HMIP returned no parseable
period. NOTE: snapshot tables can carry per-row period
variance — HMIP returns each geo's most recent release, and
sparser geos may sit on older data.
sub_geography str | null Provinces / Centres / CSD / etc. breakdowns. Null when the
table is queried at the geo itself (no sub-geos to list).
category str dimension value: 'Apartment', 'Studio', '$1,000 - $1,249', …
value f64 | null suppressed → null
reliability str | null 'a' excellent → 'd' poor
survey, table_id, geography metadata; geography = the geo we queried
| Query | Result | Matches reality |
|---|---|---|
| Canada apartment starts 2025 (annual, summed from 5.7.2 quarterly) | 94,796 | ✓ |
| Canada apartment starts 2018 → 2025 trend | 63,771 → 94,796 | ✓ |
| Vancouver CMA total vacancy 2025 | 3.7% reliability a |
✓ |
| Provincial vacancy ranking 2025 | Alberta 4.2% > BC 3.5% > ... | ✓ |
| March 2026 provincial starts | Alberta 1,169 leads | ✓ |
| Toronto total vacancy 2023 → 2025 | 1.4 → 2.5 → 3.0 | ✓ |
Every bulk_pull writes a JSONL manifest to data/logs/{label}_{utc_timestamp}.jsonl. One record per non-skipped attempt:
{"ts":"2026-05-23T03:04:27Z","survey":"Rms","table_id":"2.1.3.1","geography":"Canada",
"outcome":"ok","latency_s":0.233,"error_class":null,"error_msg":null}Cross-run analysis via DuckDB:
-- Most-frequent errors by table
SELECT table_id, count(*) AS n
FROM 'data/logs/*.jsonl'
WHERE outcome = 'error' GROUP BY 1 ORDER BY 2 DESC;
-- Latency distribution
SELECT outcome, count(*) AS n, avg(latency_s) AS avg_s, max(latency_s) AS max_s
FROM 'data/logs/*.jsonl' GROUP BY 1;Discovery write-ups live in DATA_DISCOVERY.md — append-only log of catalogue / tidy bugs, with scripts/probe_table.py as the diagnostic tool. Recent finds: slash vs hyphen CSD-name drift (2026-06-10, 5 CSDs recovered); Srms publishes for 8 Ontario CMAs not 4 (2026-06-09, 4 CMAs recovered); RMS bedroom-filter bug (2026-05-23, 9 dimensions recovered); tidy() snapshot/single-geo row loss (2026-05-23).
| # | Issue | Severity | Notes |
|---|---|---|---|
| 1 | 2 endpoints return persistent 500s at every geography (1.16.3.4, 1.16.3.5) |
Low | Stale R catalogue entries; HMIP-side bug. Denylisted in validity.BROKEN_TABLE_IDS. |
| 1b | 12 Scss tables 500 for a specific set of 8 CMAs (per-province) | Low–Medium | See "HMIP 500s on specific (table, CMA) pairs" below. NOT denylisted — they work for ~35 other CMAs each. |
| 2 | RESOLVED for Ontario CSD + CT | Ported via extract_geo_lookups.py. Neighbourhoods + Survey Zones derivable from the CT crosswalk (columns present in cts_ontario.csv); not yet exposed as separate Geography sets. |
|
| 3 | French (/fr/) endpoints not pulled |
Low | Possibly different metadata, possibly redundant. Defer. |
| 4 | No Open Government Portal sweep | Medium | ~37 CMHC datasets there in CSV/XML, easy to scrape, gives national-level cross-check data. |
| 5 | RESOLVED | _tidy_snapshot() fallback. |
|
| 6 | RESOLVED 2026-05-23 | Dropped from _RMS_SERIES. Recovered Rent Ranges / Rent Quartiles / Year of Construction etc. — see DATA_DISCOVERY.md. |
|
| 7 | tidy() losing snapshot period + single-geo rows |
RESOLVED 2026-05-23 | Added subtitle period extraction; preserve empty-index rows when no other rows exist. Recovered ~30k previously-silenced parquet rows. |
| 8 | RESOLVED 2026-06-09 | pull_csds.py --refresh-empty-days 0 rerun; CSD-level Rms now expanded from ~140k rows to ~440k. CT pull since completed (item 10). |
|
| 9 | RESOLVED 2026-06-10 | Added cmhc.geographies.normalize_name(). Mart now matches both forms. |
|
| 10 | RESOLVED 2026-06-17 | All 15 tract-publishing Rms tables pulled across 2,382 CTs (after the 2026-06-16 leaf-redundancy guard + the 2026-06-17 exclusion of 11 CT-500 tables). 1,589 CTs returned data; parquet + mart rebuilt. Biggest remaining Ontario-rental coverage gap, now closed. |
First full Ontario CMA pull (2026-05-22) surfaced 96 deterministic HTTP 500s that survived the 3-retry/exponential-backoff:
- 12 Scss tables affected (all
breakdown=Historical Time Periods,geo_filter=Default):1.2.1, 1.2.2, 1.2.3, 1.2.4, 1.2.5, 1.2.6, 1.2.7, 1.2.81.9.31.16.1, 1.16.2, 1.16.5
- Same 8 CMAs error across all 12 tables — alphabetically the first 8 Ontario CMAs:
Barrie, Belleville – Quinte West, Brantford, Guelph, Hamilton, Kingston, Kitchener – Cambridge – Waterloo, London
- Same 12 tables succeed (
ok) for the other ~35 Ontario CMAs — so the tables themselves work. - Cause unknown. Hypotheses: HMIP backend has per-(table, CMA) data gaps that error instead of returning empty; or these CMAs trigger a bug in HMIP's table-generation code. Worth re-checking on a future provincial pull to see if the same pattern holds in other provinces (which would suggest a position-in-batch artifact rather than CMA-specific data gaps).
We deliberately do not denylist these — denylisting at the table_id level would block ~315 valid datasets (35 working CMAs × 9 tables). A pair-level denylist (set[(table_id, geo_name)]) is the right tool if the noise becomes painful; deferred until we see whether the pattern is stable across provinces.
- Refresh Census / Seniors / Core Housing Need at Ontario CMA with
--refresh-empty-days 0. The 2026-06-09 Srms recovery showed pre-fix empty markers were hiding 4 of 8 publishing CMAs; the same drift likely applies to these surveys. - Pull remaining provinces' CMAs —
pull_cmas.py --province NAMEfor BC, Alberta, Quebec, etc. Each ~10 min at concurrency=5. Currently optional; widen scope only on explicit ask.
- Static data tables — separate harvest, shared schema. Second data source, structurally separate from HMIP at acquisition, converging at the schema. Full design in PLAN.md. Pipeline is built end-to-end (catalogue → download → matrix engine → parquet); 18 tables spec'd, ~19k rows (see "Static data tables" under Current data). Remaining:
- Spec the rest of the high-value uniques — mortgage-and-debt (credit, debt obligations, arrears) and rental-market uniques (seniors survey, non-resident ownership, percentile rents). Each is a screened
specs.pyentry;build_static_parquet.pypicks them up automatically. - Multi-dimension engine mode — unlock the ~20 deferred household-characteristics tables with a leading Tenure / Age group / Quintile dimension. The biggest single coverage gain in the static surface.
- Structural stragglers — the 15 no-header / by-geography household-characteristics files; individual handling.
- Source-agnostic mart builder — union
data/clean/(HMIP) +data/clean/static/, reconcile geography vianormalize_name()(static names aren't internally consistent —NewfoundlandvsNewfoundland and Labrador).
- Spec the rest of the high-value uniques — mortgage-and-debt (credit, debt obligations, arrears) and rental-market uniques (seniors survey, non-resident ownership, percentile rents). Each is a screened
- Catalogue probe sweep — use
scripts/probe_table.pyagainst other dimensions with non-trivial filter sets. The bedroom-filter and slash/hyphen bugs almost certainly aren't the last stale entries. - Open Government Portal sweep — separate
pull_opengov.py. Direct CSV downloads, gives independent cross-check data. - Pair-level denylist if the per-(table, CMA) 500 noise becomes painful across provinces (see issue 1b).
- Srms sub-CMA validity filter — restrict to the 8 publishing Ontario CMAs (Barrie, Hamilton, Kitchener – Cambridge – Waterloo, London, Ottawa, St. Catharines – Niagara, Toronto, Windsor). Mirrors the 2026-05-23 non-CMA-CSD optimization. See DATA_DISCOVERY.md for context.
- Geography vintage tagging — record census year per row to flag boundary changes (matters more once we have CT data spanning multiple census vintages).
- StatCan-sourced concordance catalogue — vintage-tagged CSD → CMA / CD / Province crosswalk from StatCan SGC/GAF. Would replace the current ad-hoc lookups and let the mart populate
cmafor Census-Agglomeration CSDs that currently havecma=NULL. Multi-week effort; deferred. - Neighbourhoods + Survey Zones as first-class Geography sets (data already in
cts_ontario.csvvia NBHDCODE / ZONECODE columns). - Other-domain data marts — Scss (housing starts), Census, Core Housing Need each warrant their own DuckDB mart following the rental mart's pattern. On demand.
- Static publications crawler — only if HMIP + Open Gov together leave gaps worth filling.
- ✅ Ontario CT pull complete (2026-06-17): all 15 tract-publishing Rms tables across 2,382 CTs; 1,589 CTs returned data (~1.29M CT rows). Parquet +
cmhc_rental.duckdbmart rebuilt to CT level. Closes open item #10 and the largest Ontario-rental coverage gap. - ✅ Massive CSD-level Rms recovery (2026-06-09): re-pulled with
--refresh-empty-days 0, +9,326 new (table, CSD) combos returning data, +461k rows in the parquet. - ✅ Ontario rental data mart shipped (2026-06-10):
data/marts/cmhc_rental.duckdb. See DATAMART.md. - ✅ Shiny for Python choropleth app (
app/shiny/) — three pages live: CSD rent map, CMA rent map, charts. - ✅ Srms 4 → 8 Ontario CMA recovery (2026-06-09).
- ✅ Slash/hyphen CSD name normalization (2026-06-10) —
cmhc.geographies.normalize_name.
See PLAN.md "Things we are deliberately not doing."