An antenatal preterm-birth risk model on CDC NVSS Natality data — built and reported the way a clinical prediction study should be. Strict early-pregnancy feature discipline (an enforced leakage guard), temporal external validation on non-overlapping birth years, full TRIPOD+AI (2024) reporting, SHAP explainability, calibration and decision-curve analysis against a clinical comparator, and race/ethnicity + maternal-age subgroup fairness — all reproducible with one command.
About this repository. This is a purpose-built methodological sample by Moocean Studio. To keep it fully reproducible offline it ships a schema-accurate synthetic generator that mirrors the real CDC NVSS Natality public-use file (same columns, code values and plausible joint structure); the reported numbers come from that synthetic data and are illustrative. Pointing the pipeline at the real, open CDC files is a documented one-liner (
scripts/download_nvss.md) and changes no downstream code. The point of the repo is to demonstrate the method and reporting rigour end to end.
Most published work on the CDC natality files is descriptive — rates and trends. A score that is actually useful has to predict, early in pregnancy, which pregnancies are heading toward preterm birth — explainably, and with performance checked across subgroups. Two disciplines make or break that:
- No leakage. If the feature set includes gestational age, birth weight,
Apgar, NICU, or any labour/delivery variable, the endpoint becomes trivially
predictable and the model is clinically useless. Here, a
LEAKAGE_BLOCKLISTenumerates every such field and anassert_no_leakage()guard runs before every model fit — leakage raises an error instead of quietly inflating AUC. - Real external validity. We develop on earlier birth years (2015–2017) and validate on a later, non-overlapping national cohort (2019–2020) — a genuine temporal holdout under distribution shift, not a random split of one year.
Endpoint: preterm birth (<37 weeks). Development n=120,000 (9.8% preterm); external validation n=80,000 (10.1%). Selected model: histogram gradient boosting, isotonic-calibrated.
| Metric | Development (apparent) | Internal 5-fold CV | External (temporal) |
|---|---|---|---|
| AUC | 0.691 | 0.649 | 0.658 (95% CI 0.651–0.665) |
| Brier | 0.083 | — | 0.086 |
| Calibration slope | 1.10 | — | 1.06 |
| Calibration-in-the-large | +0.13 | — | +0.16 |
The close agreement between internal CV (0.649) and external (0.658) AUC shows the internal estimate generalised; the modest apparent→external drop is the expected signature of optimism plus temporal drift. AUC ≈0.66 is realistic — population preterm prediction from routine antenatal factors is genuinely hard, and this is an honest number, not a hero metric.
A fairness finding worth highlighting. Discrimination is similar across groups,
but calibration is not: under-prediction is largest in the smallest subgroups
(mothers ≥40, "other" race/ethnicity). Equal discrimination with unequal
calibration is a well-known fairness failure mode — a single global threshold would
mis-target treatment across groups. The pipeline surfaces this instead of hiding it
behind a pooled metric. Full per-subgroup table: reports/model_card.md.
make setup # create venv, install deps
make test # 22 tests: leakage guard, schema, pipeline, metrics
make run # regenerate every figure + reports/results.json + model_card.mdOr directly:
python scripts/run_pipeline.py --endpoint preterm --profile early --n-per-year 40000
python scripts/run_pipeline.py --endpoint nicu # alternative endpoint
python scripts/run_pipeline.py --profile anytime # wider feature windowRun on the real CDC files: scripts/download_nvss.md.
src/perinatal/
schema.py Antenatal feature whitelist, leakage blocklist + guard, 2003
birth-certificate revision handling, endpoint definitions.
data_loader.py Real fixed-width PUF parser + schema-accurate synthetic loader;
label derivation kept separate from features (no leakage).
synth.py Causal-ish synthetic generator with realistic risk structure,
built-in racial/socioeconomic disparities, and temporal drift.
features.py Feature engineering + 'early'/'anytime' profiles; calls the
leakage guard before returning the design matrix.
model.py Logistic regression vs gradient boosting, isotonic calibration.
evaluate.py AUC + bootstrap CI, calibration curve/slope, decision-curve
analysis vs a clinical comparator, subgroup metrics.
explain.py SHAP (tree) with permutation-importance / coefficient fallback.
plots.py Publication-style figures (matplotlib only).
pipeline.py End-to-end: cohorts → features → CV select → validate → report.
Field availability changed when the U.S. Standard Certificate was revised in 2003
(states adopted it through 2014). schema.variables_available_for_years() drops
revision-specific fields (BMI, WIC, checkbox risk factors, cigarettes-per-trimester)
for any analysis reaching before 2014, so a model is never trained on a feature
that doesn't exist for part of its cohort. The default cohorts stay in the fully
revised era.
Every reporting element the guideline asks for is implemented and mapped
item-by-item in manuscript/TRIPOD-AI-checklist.md:
data leakage prevention (enforced), calibration plots (not just slope/intercept
in text), decision-curve analysis against a clinical comparator, subgroup fairness
separating discrimination from calibration, and SHAP explainability. A complete
manuscript draft in the required structure (Abstract ≤300 words, Keywords,
Introduction, Methods, Results, Discussion, Conclusion, Vancouver references) is in
manuscript/manuscript.md, with a dated
prior-work/novelty search template in
manuscript/literature-search.md.
- Numbers are from synthetic data and are illustrative; the substantive finding requires the real CDC files.
- Race/ethnicity is used as a fairness axis and social-risk proxy, not a biological cause of preterm birth.
- Birth-certificate risk factors are under-reported and vary by reporting area; real-world performance will be attenuated. This is a population-level triage aid, not for clinical use.
MIT — see LICENSE.





