A global registry of data portals, catalogs, data repositories, and related data infrastructure.
Working tree (12 September 2026): 37,170 verified catalogs · 486 software platforms · 224 countries and territories · 0 scheduled YAML records.
Last published snapshot: v1.21.0, 12 September 2026 (37,170 catalogs · 486 software · 0 scheduled).
This is the catalog-metadata pillar of the Common Data Index / open search engine. It describes catalogs (open data portals, geoportals, scientific repositories, indicator sites, and similar infrastructure), not the datasets those catalogs hold.
- Source of truth: YAML under
data/entities/ - Consume: JSONL, Parquet, and DuckDB under
data/datasets/— do not parse thousands of YAML files - Out of scope: this repository does not host a production query API or MCP server
Code is MIT; data and documentation are CC BY 4.0. Inspired by re3data and FAIRsharing, with a broader focus on open data of every kind — government, geospatial, scientific, and statistical — not only research data.
Python 3.10–3.12. Install deps with pip install -r requirements.txt. Nested fields in DuckDB and Parquet are native STRUCT / LIST types; query them with field access, not LIKE on JSON text.
duckdb data/datasets/datasets.duckdb \
-c "SELECT id, name, link FROM catalogs WHERE software.id = 'ckan' LIMIT 10;"import duckdb
con = duckdb.connect("data/datasets/datasets.duckdb")
con.execute(
"""
SELECT id, name, link
FROM catalogs
WHERE catalog_type = 'Open data portal'
AND list_contains(
list_transform(coverage, x -> x.location.country.id),
'US'
)
LIMIT 10
"""
).fetchall()More patterns: docs/query-examples.md. Join keys and column types: docs/ai-consumers.md.
Last published snapshot (v1.21.0, 2026-09-12: 37,170 catalogs, 486 software, 0 scheduled). Working-tree exports below match source YAML. Record-count contract: docs/exports.md.
| File | Contents |
|---|---|
data/datasets/catalogs.jsonl.zst |
37,170 verified catalog records |
data/datasets/software.jsonl (+ .zst) |
486 software / platform definitions |
data/datasets/scheduled.jsonl (+ .zst) |
0 scheduled sources (0 YAML files) |
data/datasets/full.jsonl (+ .zst) |
Entities + scheduled (37,170) |
data/datasets/full.parquet, data/datasets/datasets.duckdb |
Analytics-friendly copies of full.jsonl |
Rebuild from YAML (never hand-edit data/datasets/):
python scripts/builder.py buildDecompress .zst with unzstd file.zst. Filter by catalog_type or software.id in DuckDB or Parquet; there are no pre-sliced bytype/ or bysoftware/ dumps.
Each record is one catalog: name, URL, owner, geographic coverage, software platform, API/harvest endpoints, and optional identifiers (Wikidata, re3data, OpenAIRE, …).
catalog_type |
Folder | Typical contents |
|---|---|---|
| Open data portal | opendata/ |
Government and institutional open data |
| Geoportal | geo/ |
Spatial data, OGC services, map viewers |
| Scientific data repository | scientific/ |
Research data, CRIS, institutional repos |
| Indicators catalog | indicators/ |
Statistical indicators, SDMX, dashboards |
| Microdata catalog | microdata/ |
Survey / census microdata |
| Machine learning catalog | ml/ |
ML datasets and models |
| Data search engine | search/ |
Cross-catalog search / aggregators |
| API Catalog | api/ |
API directories |
| Data marketplace | marketplace/ |
Commercial data markets |
| Metadata catalog | metadata/ |
Metadata registries |
| Other | other/ |
Uncategorized |
Use this registry to find portals by country, type, or software, join catalogs to external identifiers, or feed a downstream harvester from endpoints[]. Do not use it to search for a dataset by title — harvest the remote catalog instead (docs/harvest.md). Scope: docs/when-to-use.md.
Published internals for humans and coding agents:
| Goal | Start here |
|---|---|
| Docs site | https://datenoio.github.io/dataportals-registry/ |
| Getting started | docs/getting-started.md |
| Field reference and vocabularies | docs/data-model.md, docs/vocabularies.md, docs/catalog-types.md |
| Software IDs | docs/software-index.md, docs/software-taxonomy.md |
| Find catalogs not yet registered | docs/discovery.md |
| Harvest datasets from a catalog API | docs/harvest.md |
| CLI | docs/cli.md |
| Agent index | llms.txt (also /dataportals-registry/llms.txt on the docs site) |
Source markdown lives in docs/. Local preview: cd website && npm install && npm run start. Working notes stay in devdocs/ and are not on the site.
Catalog YAML: data/entities/{COUNTRY}/{Federal|SUBREGION}/{type}/{id}.yaml. Filename must equal id (lowercase letters and digits only). Layout and field rules: docs/directory-layout.md, docs/data-model.md.
data/entities/{CC}/{Federal|SUBREGION}/{type}/{id}.yaml verified catalogs
data/scheduled/ unverified (promote later)
data/software/ platform definitions
data/schemes/ Cerberus + JSON Schema
data/reference/ controlled vocabularies
data/datasets/ generated exports (do not edit)
Example — FAA Open Data Portal (data/entities/US/Federal/opendata/catalogdatafaagov.yaml):
id: catalogdatafaagov
uid: cdi00005263
name: Federal Aviation Administration Open Data Portal
link: https://catalog.data.faa.gov
catalog_type: Open data portal
access_mode:
- open
status: active
api: true
api_status: active
software:
id: ckan
name: CKAN
owner:
name: Federal Aviation Administration
type: Central government
location:
country:
id: US
name: United States
level: 20
coverage:
- location:
country:
id: US
name: United States
level: 20properties.is_national: true only for the country’s official catalog of that type (national open-data portal, NSDI/geoportal, or NSO product) — not because the owner is a federal agency. See docs/data-model.md.
Already in this registry
- By geography:
data/entities/{COUNTRY_CODE}/(for exampleUS,FR,BR).Federal/is national/central; subregion folders are ISO 3166-2 style (US-CA,GB-SCT,BR-SP). - By type: under each country,
opendata/,geo/,scientific/,indicators/,microdata/,ml/,search/,api/,marketplace/,metadata/,other/. - By software: filter
software.idin DuckDB (canonical IDs: docs/software-index.md). Platform YAML indata/software/also hascategoryandsubtypefor self-hosted vs SaaS vs protocol-first comparisons. - By URL / id: query
catalogsrather than walking YAML:
SELECT id, name, link, catalog_type, status
FROM catalogs
WHERE lower(link) LIKE '%data.faa.gov%'
OR id = 'catalogdatafaagov';Not yet in this registry
Vendor galleries, search-engine recipes, and per-platform fingerprints: docs/discovery.md (humans) and docs/agents/discover.md (agents). Search tools: docs/discovery-search-tools.md. LLM / MCP setup: docs/discovery-agent-tools.md. Endpoint fill: docs/apidetect.md. URL liveness: docs/liveness.md.
Quality analysis flags duplicate URLs, path/owner country mismatches, non-canonical owner types, schema violations, and related issues. CI guards regressions with dataquality/baseline_counts.json. Issue codes: docs/quality-rules.md.
python scripts/builder.py validate-yaml
python scripts/builder.py analyze-qualityReports land in dataquality/ (full_report.txt, primary_priority.jsonl, plus per-country and per-priority breakouts). Helper scripts scripts/fix_*_issues.py apply automated fixes by priority. Workflow: docs/metadata-quality.md.
Fixes and new catalogs are welcome via pull request or issue. Full guide: CONTRIBUTING.md. Agents: docs/agents/contribute.md. What to hunt next: docs/agents/improve.md.
python scripts/builder.py add-single "https://example.com/data" \
--software ckan \
--catalog-type "Open data portal" \
--name "Example Data Portal" \
--country US \
--scheduled
python scripts/builder.py assign
python scripts/builder.py validate-yaml --id examplecomPrefer --scheduled for unverified finds; promote later (docs/scheduled.md). Duplicate link values fail quality checks.
| Pipeline | Script | Docs |
|---|---|---|
Re3Data metadata into _re3data |
python scripts/re3data_enrichment.py enrich --dry-run |
docs/re3data.md |
| CKAN sites from ecosystem.ckan.org | python scripts/sync_ckan_ecosystem.py --dry-run |
docs/ckan-sync.md |
| OpenAIRE Graph data sources | python scripts/extract_openaire_portals.py list-sources --output /tmp/openaire_sources.json |
docs/openaire-sync.md |
See CITATION.cff and DATASHEET.md (purpose, bias, limitations).
dataportals-registry: A global registry of open data portals and catalogs
(Common Data Index, 2026). CC-BY-4.0.
https://github.com/datenoio/dataportals-registry
- Code: MIT
- Data: CC BY 4.0
- SECURITY.md — vulnerability reporting
- CODE_OF_CONDUCT.md — community standards