Submission for IEMH4-HC-01. "Swasthya" (health) + "Setu" (bridge) — the bridge between a village and the health system.
An offline-first PWA for ASHA/ANM workers: enter a patient's vitals (typed or by voice, in any of 10 Indian languages) → get an explainable diabetes risk score and a non-invasive photo-based anemia band → dangerous cases are caught by a safety guardrail and triaged to the nearest PHC → an officer dashboard shows every case on a village outbreak radar. Works with the network off and syncs when it returns. Exports FHIR-style JSON.
Every result is a screening aid, not a diagnosis. That string is on every risk result in the UI, in every API response, and in the FHIR bundle.
Two terminals. Requires Python 3.11+ (3.12 used here) and Node 20+.
Backend
cd backend && python -m venv venv && venv\Scripts\activate
pip install -r requirements.txt
python ml/train_diabetes.py
set HC01_API_KEY=hc01-dev-key && uvicorn app.main:app --reload --port 8000Frontend
cd frontend && npm install
npm run devOpen http://localhost:5173. The database is created and seeded with the demo story on first backend start.
The API key is not optional. Every route that can reach patient data requires an
X-API-Keyheader, and the backend fails closed — start it withoutHC01_API_KEYand it generates a random one intobackend/.hc01_api_key, at which point the frontend's committed dev placeholder no longer matches and every request returns 401. Either exportHC01_API_KEY=hc01-dev-keyas above, orcat backend/.hc01_api_keyintofrontend/.env.localasVITE_API_KEY=…. See Access control.
Port note for this machine:
8000is already held by another process (Manager), sofrontend/.env.localpoints the app at8001. Start the backend with--port 8001, or delete.env.localand use the default 8000 once that port is free.
Verify
cd backend && venv\Scripts\python -m pytest tests -q58 tests, one per acceptance criterion in the spec plus the access-control
suite. python tests/bench.py 100000 runs the scale benchmark.
cd frontend && npm run check:i18nFails if any of the 10 locales drifts from en.json. i18n falls back to English
for a missing key, so drift is invisible at runtime — this is what makes it
visible at build time.
| Capability | Status |
|---|---|
| Explainable diabetes risk (LogisticRegression + exact linear attribution) | built |
| Offline-first PWA intake + instant result (IndexedDB queue, Workbox SW) | built |
| Voice input (Web Speech) + i18n in 10 of the 23 official languages | built |
| Access control: shared API key, rate limiting, security headers | built |
| Photo anemia screening (OpenCV pallor pipeline) | built |
| Encrypted health-record storage (AES-256-GCM at rest + blind index) | built |
| Regional-language assistant: FAQ + pre-diagnosis triage | built, offline |
| Safety guardrail: abstain-when-uncertain + emergency red-flags | built |
| Village outbreak radar (Leaflet heat + village clustering) | built |
| Worker/officer dashboard: paged list, filters, 14-day trend | built |
| Nearest-PHC lookup (haversine) + auto-flag | built |
FHIR-style export (Bundle / Patient / Observation / RiskAssessment) |
built, static shape |
| Teleconsult | stub — real booking record + Jitsi room URL, no signalling |
| ABDM/NDHM interop | stub — FHIR-shaped JSON, no registry integration |
| OCR / SMS-IVR | not built (roadmap) |
Client (PWA) Server (FastAPI)
───────────────────────────── ────────────────────────────────
Intake ── voice ── camera /predict/diabetes ─┐
│ /predict/anemia-photo │
▼ /patients /sync ├─ guardrail
IndexedDB (patients + queue) /patients/{id}/fhir │ (always)
│ /phcs /outbreak ─┘
▼ online / 'online' event │
syncEngine.flush() ──── POST /sync ───────► SQLite (WAL)
│
Dashboard ── Leaflet heat ── Recharts ◄──────────┘
The client is authoritative offline; the server is authoritative once synced.
Sync is last-write-wins keyed by client_uuid, so retries are free and a record
is never recorded twice.
backend/
app/
main.py FastAPI app, CORS, gzip, lifespan (init + model warm)
config.py every threshold and limit, env-overridable
db.py engine, SQLite pragmas (WAL), init_db, one-column migration
models.py SQLModel tables + the indexes the three query shapes need
schemas.py request/response contracts
errors.py the single error envelope + global handlers
deps.py session + capped pagination dependencies
seed.py the 16-patient demo story
routers/ predict.py patients.py phcs.py
services/ diabetes.py anemia_vision.py guardrail.py
outbreak.py fhir.py crypto.py
ml/
train_diabetes.py deterministic training -> .pkl + feature_meta.json
calibrate_anemia.py re-fit the pallor thresholds on labelled images
tests/
test_acceptance.py the spec's definition of done, executable
bench.py scale benchmark
frontend/src/
api/client.js axios + one normalized error shape
db/idb.js IndexedDB stores, cursor paging
sync/syncEngine.js single-flight, confirmed-before-delete queue flush
lib/guardrail.js offline mirror of the server guardrail
lib/voiceParse.js speech -> intake fields (with a dev self-check)
lib/chatbot.js FAQ + triage routing, reuses guardrail + symptom vocabulary
lib/geo.js client-side haversine so "nearest PHC" works offline
pages/ Intake Result Assistant Dashboard
components/ VoiceButton CameraCapture RiskCard RedFlagBanner
OutbreakMap TrendChart ErrorBoundary
Pima Indians Diabetes, restricted to the five features an ASHA worker can
actually collect — pregnancies, glucose, blood_pressure, bmi, age — plus
family_history_diabetes, engineered from DiabetesPedigreeFunction > median.
Physiologically impossible zeros in glucose / BP / BMI are median-imputed.
LogisticRegression(max_iter=1000, random_state=42, class_weight="balanced"),
test_size=0.2, everything seeded at 42. Current holdout: accuracy 0.708,
ROC-AUC 0.803. Identical input always produces identical output — asserted by
test_determinism_same_input_same_output.
Explanation is exact, not approximated. For a linear model,
contribution_i = coef_i × standardized_value_i, and the contributions plus the
intercept are the log-odds. The top three by absolute contribution are
returned, normalized so their magnitudes sum to 1. No SHAP required — which is
also why shap is not in requirements.txt.
Runs after inference, before any result leaves the server, on every path
(/predict/diabetes, /patients, /sync). Two independent layers:
A. Abstain when uncertain — a vital outside the spec's clinical range, or a
probability in [0.45, 0.55], returns band: "uncertain" and
message_key: "result.need_recheck" with HTTP 200. Never a 500, never a
confident guess.
B. Emergency red-flags — pure rules that run ahead of the score and are never suppressed by it:
| Rule | Condition |
|---|---|
cardiac |
chest_pain AND (sweating OR breathlessness) |
respiratory_distress |
breathlessness AND fever |
hyperglycemic_emergency |
glucose >= 300 |
severe_hypertension |
blood_pressure >= 120 (diastolic) |
severe_anemia |
anemia_band == "severe" |
Which rule fired is logged and returned. The seeded patient Bhola Nath has a
7% diabetes risk and still escalates — that is the point of the layer, and
test_chest_pain_and_sweating_red_flags_even_when_risk_is_low locks it in.
frontend/src/lib/guardrail.js is a byte-for-byte mirror of these rules so an
offline worker gets the same escalation. Both have self-checks.
pallor_index = 1 - (R - G) / R, clamped to [0,1], over a region of interest:
the lower half of a Haar-detected eye when a face is in frame, otherwise the
spec's centre-lower crop.
- Foreground masking — pixels below luma 25 are background (deep shadow, or the black surround of a segmented crop), then the darkest tenth (eyelash) and brightest 3% (specular) of the remaining tissue are trimmed.
- Abstain on bad input — ROI too dark or too small returns
band: "unknown",abstain: true. Corrupt or oversized files returnIMAGE_ERROR. Neither crashes.
Measured against CP-AnemiC (710 labelled paediatric conjunctiva images with haemoglobin values, Mendeley doi:10.17632/m53vz6b7fx.1):
| Comparison | AUC |
|---|---|
| anemic vs non-anemic | 0.51 |
| severe (Hb<7) vs normal (Hb≥11) | 0.53 |
Median index by true severity: normal 0.621 · mild 0.628 · moderate 0.637 · severe 0.600. Flat — and severe scores lower than normal. This is a coin flip. No variant tested beat AUC 0.57 (normalised difference, CIELAB a*, HSV saturation, erythema index, reddest-quartile, with and without white balance). Gray-world normalisation is specifically not used: it forces R≡G and drives this index to exactly 1.0 for every image.
Mean conjunctival colour does not identify anemia in that cohort. Published work that reaches useful accuracy here uses CNNs, not colour statistics.
What that changed in the code:
- Thresholds recalibrated from
0.30/0.45/0.60(spec placeholders, which put 93% of real photos in moderate-or-severe) to0.60/0.65/0.76, the quantiles that reproduce the cohort's WHO severity prevalence. Severe went from 48% of images to 8.5%, against a true prevalence of 6.8%. - Foreground masking replaced percentile trimming. The old rule left the
black surround in the average, so mean ROI brightness read 30 and the
pipeline abstained on 74% of valid clinical photos. Now 0%. Locked in by
test_segmented_crop_on_black_background_is_not_mistaken_for_darkness. - The measured AUC is exposed at
GET /meta, so the claim is auditable rather than folklore. calibrate_anemia.pynow reports AUC first and prints a warning when a pair is not separable — fitting cut points to a set the index cannot separate redistributes labels without making them predictive.
So: this ships as a pallor measurement with an honest caveat, not a
validated anemia screen. It is a prompt to look at the patient. The spec's
severe_anemia red-flag rule is kept as written — flagged as a concern, since
an unvalidated band driving an emergency escalation will produce false
positives at roughly its base rate.
To re-fit for a different camera or population:
cd backend && venv\Scripts\python ml/calibrate_anemia.py samples/with samples/normal|mild|moderate|severe/. Read the AUC line first.
Patient records are sealed with AES-256-GCM before they touch storage. Name,
age, sex, pregnancies, glucose, blood pressure, BMI, family history and symptoms
live inside one authenticated ciphertext per row, bound to the record's
client_uuid as associated data — so a row copied over another one fails
authentication instead of silently impersonating it.
What stays plaintext, and why. Village, coordinates, timestamps, risk band
and flags remain queryable: you cannot GROUP BY a ciphertext, and the outbreak
radar and dashboard exist. What leaks without the key is pseudonymous —
"someone in Rampur, high risk, last Tuesday" — never a name or a vital.
Name search survives via a blind index: an HMAC of each lowercased name
token, so ?q=devi still works while the server never stores a readable name.
Two things the migration had to scrub. Encrypting the table is not enough,
and grepping the file proved it: after migrating 21 records, every name was
still readable on disk. They were sitting in the write-ahead log and on the
freelist that DROP TABLE left behind. The migration now checkpoints the WAL to
zero and VACUUMs the file, and PRAGMA secure_delete=ON stops future deletes
leaving readable residue. Verified by scanning hc01.db, -wal and -shm as
raw bytes:
Sunita Devi readable on disk: False Rampur (village) still queryable: True
Bhola Nath readable on disk: False
test_patient_identifiers_are_not_stored_in_plaintext does exactly that scan on
every run, so this cannot silently regress. Tampering is covered too — flipping
one ciphertext byte makes that record fail loudly, while the dashboard reports
unreadable: 1 and keeps rendering the rest.
Key management. HC01_ENCRYPTION_KEY (base64, 32 bytes) in any real
deployment. Absent it, a key is generated once into backend/.hc01_key
(gitignored) with a loud warning — the app never silently falls back to
plaintext. Lose that file and the records are unrecoverable.
Threat model, stated plainly: this defeats a stolen laptop, a copied database, or a leaked backup. It does not defeat someone who can call the API — the server decrypts for any caller it accepts, by design. That is a different control, and it is the next section.
Encryption at rest and access control answer different attacks, and shipping one
without the other is how a system ends up "encrypted" and wide open at the same
time. Before this layer, GET /patients/1 returned a complete medical record to
anyone who could reach the port — and ids are sequential, so the whole database
enumerated in a single loop.
Every route that can reach patient data now requires X-API-Key. One
dependency (app/security.py), applied to all three routers:
_protected = [Depends(require_api_key)]
app.include_router(patients.router, dependencies=_protected)/health and /meta stay open — they carry thresholds and a version string, and
making a liveness probe present a PHI credential is a bad trade. The comparison
is hmac.compare_digest, not ==: string equality short-circuits on the first
differing byte and leaks the shared prefix to a timing attacker.
It fails closed. No HC01_API_KEY in the environment means a random key is
generated into backend/.hc01_api_key (gitignored) and logged loudly — an unset
key never means "allow everyone", which is precisely the bug this exists to fix.
What it buys, and what it does not. The key ships inside the frontend bundle,
so it authenticates the deployment, not the worker: internet-wide scanning and
id enumeration are closed; anyone holding the app still holds the key, there is
no per-worker audit trail, and revoking one lost phone means rotating everybody.
The upgrade path is a workers table and a PIN → short-lived signed token, and
require_api_key is the single seam every PHI route already passes through, so
that swap touches one file.
Rate limiting — a 120-request/60 s sliding window per IP, returning 429
with Retry-After. Sized to be invisible to real use (a whole intake is under
ten requests; a 2,000-record sync is four chunks) while turning "enumerate every
patient id" from seconds into hours. X-Forwarded-For is deliberately not
honoured — it is caller-controlled, and trusting it would let one attacker forge
a fresh bucket per request. In-process counters, so the cap multiplies behind
multiple uvicorn workers; a shared store is the fix if it ever runs multi-process.
Security headers on every response: X-Content-Type-Options: nosniff,
X-Frame-Options: DENY, Referrer-Policy: no-referrer, Cache-Control: no-store (PHI must not sit in an intermediary cache), Strict-Transport-Security
(ignored over plain http, so it costs nothing until there is TLS), and a
default-src 'none' CSP — relaxed only for /docs, whose Swagger UI loads its
own assets and serves no patient data.
test_phi_routes_refuse_without_a_key covers all ten PHI routes, and the header
and rate-limit behaviour each have their own test. This is the part that fails
loudly if it regresses.
Ask tab — FAQ plus pre-diagnosis triage in English and Hindi, running
entirely in the browser. That is the same constraint as everything else
here: a worker needs triage advice most when the network is gone, so a
server round-trip would break the feature exactly when it matters.
Two rules keep it safe:
- Triage calls the same
redFlagRulesthe screening engine calls. The bot physically cannot contradict the risk result, because it is asking the same function. "chest pain and sweating" → emergency banner, in either language. - It never names a disease as a conclusion. It routes: emergency / get screened / self-care, and every string is an i18n key so the Hindi answer is the same answer, not a thinner one.
It shares the symptom vocabulary with the voice parser, so both understand
exactly one list. Answering triggers a handoff: tapping जाँच शुरू करें on
an emergency reply lands on /?symptoms=chest_pain,sweating with those symptoms
already ticked — the worker never re-types what they just described.
Ten FAQ topics: diabetes, prevention, anemia and iron-rich foods, blood pressure, fasting-test prep, when to go to a PHC, antenatal care, fever (malaria/dengue danger signs), data privacy, and how the app works.
- On submit the client generates
client_uuid+created_at(UTC), computes a local rule-of-thumb band and runs the full red-flag guardrail, writes to IndexedDB and enqueues — before touching the network. - The server is then tried for the real model result. Failure is not an error path the worker sees; the record is already safe.
syncEngine.flush()fires on app load, ononline, on tab focus, and on a 30 s retry loop whenever the queue is non-empty. The retry matters:navigator.onLinetracks the link, not the server, so an API outage while the browser is "online" fires no event at all and a queued record would otherwise sit there until the worker switched tabs. Flushing is single-flight, chunked at 200, and never removes a queue entry until the server has echoed thatclient_uuidback.- The service worker precaches the app shell, runtime-caches
GET /phcsand OSM tiles, and deliberately does not cache POSTs — a cached POST would double-submit a patient on reconnect. - "Nearest PHC" is computed in the browser (
src/lib/geo.js) from the unparameterisedGET /phcs. The server's?near_lat=..&near_lon=..variant exists and is tested, but only the plain URL is in the runtime cache, so sorting client-side is what makes referral work with the network down.
Verified by stopping both servers and reloading: the shell boots from the service worker, the village list resolves from the runtime cache, intake produces an offline estimate, the record queues, and the pending badge updates. Restarting only the API drains the queue ~20 s later with no user interaction and replaces both estimates with model results.
python tests/bench.py 100000 on this machine, SQLite in WAL mode:
| Endpoint | 50k rows | 100k rows |
|---|---|---|
POST /predict/diabetes |
1.3 ms | 0.9 ms |
POST /sync (1000 records) |
351 ms | 394 ms |
GET /patients?limit=100 |
9.7 ms | 5.5 ms |
GET /patients?flagged=true |
38 ms | 115 ms |
GET /patients (offset 20k) |
10 ms | 7.3 ms |
GET /outbreak |
135 ms | 223 ms |
GET /patients/trend |
56 ms | 199 ms |
GET /phcs (nearest) |
2.6 ms | 1.5 ms |
Measured with encryption on — every listed patient costs an AES-GCM decrypt. A 100-row page pays about 100 of them and still returns in single-digit milliseconds; bulk sync went from ~274 ms to ~351 ms per 1000 records. Cheap enough that turning it off would be a false economy.
What makes it hold up:
- Batch-first inference. A 2000-record sync is one
(2000, 6)matrix multiply, not 2000 Python calls. - Chunked upsert.
INSERT … ON CONFLICT DO UPDATEin chunks of 500 — no per-row SELECT-then-INSERT round trips, and idempotent by construction. - Aggregation in SQL.
/outbreakand/patients/trendareGROUP BYs; the heat layer is served pre-binned with aweightper bin and a hard point cap. - A denormalized
communicableflag so the outbreak filter is an index range scan instead of aLIKEover a comma-joined column. - Hard caps everywhere. Page size ≤ 1000, sync batch ≤ 2000, image ≤ 5 MB (read in capped chunks, so a 2 GB upload cannot exhaust memory), plus a decompression-bomb guard before OpenCV allocates anything.
- Nothing unhandled reaches the client. One error envelope, one global handler, and a React error boundary so a bad render never white-screens a worker mid-visit.
Benchmark numbers use random coordinates, which is the worst case for the heat binning; real data clusters at village centroids and groups far tighter.
Base http://localhost:8000. Interactive docs at /docs. Errors are always
{"error": {"code", "message", "field?"}} with codes VALIDATION_ERROR (422),
UNAUTHORIZED (401), NOT_FOUND (404), IMAGE_ERROR (400),
PAYLOAD_TOO_LARGE (413), RATE_LIMITED (429), MODEL_ERROR (500).
Every route below except /health and /meta requires an X-API-Key header.
| Method | Path | Notes |
|---|---|---|
| POST | /predict/diabetes |
risk + band + 3 reasons + guardrail |
| POST | /predict/anemia-photo |
multipart file + site |
| POST | /patients |
create one; returns the stored record incl. explanation |
| POST | /sync |
batch upsert by client_uuid, idempotent |
| GET | /patients |
paged; flagged, village, band, q, since |
| GET | /patients/trend |
daily buckets for the chart |
| GET | /patients/{id} · /patients/{id}/fhir |
record · FHIR bundle |
| GET | /phcs |
near_lat/near_lon → sorted by distance |
| GET | /outbreak |
heat points + village clusters + alert flags |
| POST | /teleconsult/book |
deterministic booking stub + Jitsi URL |
| GET | /health · /meta |
liveness · the thresholds actually in force |
Every threshold in /meta is env-overridable (HC01_*, see app/config.py) —
the UI reads them instead of hardcoding.
- Go offline (DevTools → Network → Offline). Enter a patient by voice in हिन्दी → instant risk and the "why".
- Snap a photo of the inner eyelid → non-invasive anemia band. No lab, no needle.
- Red-flag case — chest pain + sweating → the app escalates ahead of the score even though the risk is 7%. Then a borderline case → it abstains.
- High-risk flag → nearest PHC + a booked teleconsult slot.
- Back online → the queued record syncs itself and the offline estimate is replaced by the model result. Zoom the outbreak radar: Rampur alerts with 7 fever cases.
- Export → FHIR-style JSON. "ABDM-ready."
Seed data is built for exactly this: Sunita Devi (98% high risk), Bhola Nath (7% risk, cardiac red flag), Rajkumar (54% → abstain), and 7 Rampur fever cases so the cluster fires.
- The photo anemia band is not predictive (AUC 0.51 — see above). It is a measured pallor value with an honest caveat, and it should be described that way in the pitch. Everything around it — capture, abstain, error handling, banding, FHIR export — works.
- Live dictation needs one manual check. The full voice path is verified end
to end with an injected recogniser: locale follows the language toggle
(
en-IN/hi-IN), and a transcript fills age/glucose/BP/BMI, ticks symptoms and sets family history — in both English and Hindi, including Devanagari numerals (उम्र ६७ → 67). The only unverified link is Chrome's own audio→text step, which needs a real microphone. Press the mic once in Chrome before demoing. getUserMedianeeds a secure context — fine onlocalhost, HTTPS required to deploy. A file-input fallback covers the rest.- Teleconsult and ABDM are stubs, labelled as such above and in the pitch.
- Sync is last-write-wins; there is no field-level merge for two devices editing
the same
client_uuid. Not a scenario a single ASHA worker produces. - Authentication is one shared key per deployment, not per worker. It closes anonymous access and id enumeration (see Access control), but the key is compiled into the frontend bundle, so it cannot tell two workers apart, produces no audit trail, and cannot be revoked for one lost device. Per-worker credentials are the next real step.
- No HTTPS in the demo. Encryption at rest without encryption in transit is
half a control; the API key travels in a plaintext header over
localhosthere. HSTS is already sent, so terminating TLS in front is the whole fix. - Four dev-server advisories remain open (
npm audit: vite/esbuild/ plugin-react/vite-plugin-pwa). All four affect the development server only — the production build is static files — and clearing them requires the Vite 8 toolchain, whose peer graph does not resolve against the pinned React plugin. Deferred deliberately rather than churning the build on demo day. The exploitable ones (axios SSRF/prototype pollution, react-router open redirect) were upgraded: axios 1.19.0, react-router-dom 6.30.4, vite 5.4.21. - IndexedDB on the device is not encrypted. Meaningful client-side
encryption needs a secret the device does not already hold; a key sitting in
localStoragebeside the data is theatre. The honest upgrade is a worker PIN feeding PBKDF2 via WebCrypto, which adds an unlock step to the demo flow. - 10 of the 23 official languages ship (English, Hindi, Bengali, Marathi,
Telugu, Tamil, Gujarati, Punjabi, Kannada, Malayalam) — all 245 keys each,
enforced by
npm run check:i18n. The language menu lists only languages that have a locale file, so a name in the dropdown always does something. The 13 remaining include Bodo, Dogri, Santali and Manipuri; machine translation of clinical strings into those is not something to ship unreviewed, so they are deliberately absent rather than silently approximate. - The assistant is rule-based, not generative. Deliberate: inspectable, deterministic, works offline, and a wrong answer traces to one line. It covers ten topics and routes everything else to "ask the PHC".