This repository is a worked example of how the polio pipelines could be restructured. It takes the formA pipeline as it runs today and shows, side by side, what the same pipeline looks like once the rules below are applied.
It aims to demonstrate three things:
- A standard structure, so that every pipeline is laid out the same way and a colleague can find their way around one they have never opened.
- Shared code in one place: the base functions live in their own libraries, where they are maintained, versioned and tested once, instead of being copied into each pipeline.
- A consistent level of code quality across contributors, whether the code is written by hand or with an AI agent.
Nothing in forma/ was modified. The "after" code lives in separate directories so the two can be
read side by side:
| Directory | What it is |
|---|---|
| forma/ | The pipeline as it runs today. Unchanged. |
| forma_v2/ | The same pipeline reworked, as a worked example of the rules below. A model to copy from, not a production replacement (currently inside this repo). |
| polio_contracts/ | Library: the schema of every table the pipelines publish. Moves to its own repository (currently inside this repo). |
| polio_utils/ | Library: the code shared by every pipeline (IASO, ONA and OpenHEXA clients). Moves to its own repository. |
| tests/ | Tests belonging to the formA pipeline: its contracts, and that its output satisfies them. |
| .github/workflows/ | ci.yml lints and tests the whole repo on every pull request; deploy_forma.yml deploys one pipeline on a push to main. |
| CLAUDE.md | Context for AI agents (you can fill it with additional rules). |
In short:
- Created
polio_contracts, a library that declares the schema of every table the pipelines publish: its columns, the kind of values they hold, and the key that identifies a row. A pipeline now conforms its output to a contract instead of listing columns by hand, which is what keeps a table consistent across every pipeline that writes it. The library has its own tests and its own CI. - Created
polio_utils, which centralises the code shared by every pipeline — the IASO pyramid client, the ONA submissions reader and the OpenHEXA client. A single copy can be maintained and tested in one place, instead of drifting from pipeline to pipeline. It too has its own tests and CI. - Both belong in libraries rather than in a pipeline precisely because they are shared: a library has one owner, one version history, and can be pulled from GitHub by any pipeline that needs it.
- They are installed at run time, so a pipeline running on OpenHEXA pulls the version it depends on — see forma_v2/requirements.txt.
- Applied all of this to formA as the example.
forma_aux.pyandpipeline.pywere cleaned up: ruff findings fixed, Google docstrings added (can be other), imports made explicit, functions renamed tosnake_case, and the@forma.taskdecorators removed, since the steps run one after the other (no paralel processing). - Added a
tests/folder holding the tests that belong to the pipeline. Today they check that the formA contracts are well formed and that the pipeline's output satisfies them; tests of the transformation rules go here as those functions are extracted. A CI workflow runs ruff and these tests on every push.
Structure
Three repositories: the pipelines, and one per shared library.
polio_iaso_pipelines/ # the pipelines
├── pyproject.toml # deps, dev/lint groups, ruff and pytest config
├── README.md # these rules
├── CLAUDE.md # context for agents and newcomers
├── .github/workflows/
│ ├── ci.yml # ruff + pytest, whole repo, every pull request
│ └── deploy_forma.yml # one per pipeline: deploys it on a push to main
├── forma/ # one directory per pipeline
│ ├── README.md # what it does, parameters, inputs, outputs
│ ├── requirements.txt # the pinned libraries, installed at run time
│ ├── pipeline.py # orchestration only
│ ├── extract.py # reads: database, IASO, ONA, files ┐
│ ├── transform.py # pure functions — the tested part │->after full refactoring
│ └── load.py # writes: conform to contracts, save ┘
└── tests/
└── test_forma.py # one test file per pipeline (transform/important ones only)
polio_utils/ # library: shared clients
├── pyproject.toml # packaging, ruff, pytest — no conftest.py needed
├── .github/workflows/ci.yml
├── polio_utils/{iaso,ona,openhexa_client,files}.py
└── tests/test_{iaso,ona,openhexa_client,files}.py # one test file per module
polio_contracts/ # library: table schemas
├── pyproject.toml
├── .github/workflows/ci.yml
├── polio_contracts/{base,forma}.py # one module per pipeline
└── tests/test_contracts.py
forma_v2/ in this repository is that layout applied as far as the cleanup went: pipeline.py and
forma_aux.py are clean and documented, but the extract / transform / load split shown above
is still ahead of it — see rule 3.
Note there is no conftest.py anywhere. Making the package importable from the tests is a setting,
pythonpath under [tool.pytest.ini_options], not a file of code.
I have listed a few basic rules on which we can focus when developing for this repository.
main is always what is deployed. Nobody pushes to it directly: a change is written on its own
branch, opened as a pull request, and merged once the checks pass and somebody else has read it.
A pipeline runs on a schedule, against production data, and feeds dashboards people act on. A
mistake merged into main is not announced by an error message — it surfaces when a country's
figures look wrong, possibly days later. The pull request is the one moment where a reviewer, the
linter and the tests all see a change before it can reach a dashboard.
- One branch, one change. A small pull request gets read; a large one gets approved.
- Name the branch after the change (
fix/missing-campaign-id), so the list of open branches reads as the list of work in progress. Delete it once merged (we can use the jira ticket also). - Deploy from
main, automatically. When a pull request is merged, CI pushes the new pipeline version to OpenHEXA. Nobody deploys by hand, so what runs in production is always a reviewed, tested commit — and "which version is live?" has an answer you can look up(see: /.github/workflows/deploy_forma.yml).
Agree on one linter and formatter, and configure it in a single file everybody shares. The same rules apply on a developer machine and in CI, so "it passes on my laptop" and "it passes on the server" mean the same thing.
- If a rule does not suit the project, remove it from the configuration, with a comment saying why — one decision taken in the open, rather than the same argument re-litigated file by file.
- Silencing is the exception and explains itself. A suppression comment that stays needs a reason next to it, or nobody later knows whether it is still needed.
- Roll it out gradually. Turning a linter on across an existing codebase offers a choice between one unreviewable pull request and a permanently red CI that everyone learns to ignore. Instead, declare which parts are expected to be clean, check only those, and extend the list as each part is cleaned.
A function should do exactly one of three things: read (database, API, filesystem), transform (data in, data out, nothing else), or write (save a file, publish a table).
A function that does two of them can only run when the other is available. A transformation welded to a database query needs a database in order to run at all — so nobody runs it in isolation, and the decision buried inside it is never checked by anything.
- The rule of thumb: when you want to test a function that reads or writes, what you really want to test is the transformation hidden inside it. Take that out first.
- A transformation takes no clock, no connection and no logger. Pass the date in as an argument; a function that reads the clock itself answers differently every day and cannot be pinned by a test.
- How you group them is a convention. Three named groups of functions in one module is enough at small scale; three modules is the natural next step, and by then the move is mechanical.
Every pipeline has a test suite, and CI runs it on every pull request. The goal is not coverage: test the functions that decide something, and only those.
| Test it | Do not test it |
|---|---|
| Filtering, deduplication, matching rules, date windows, computed columns, aggregations | Functions that only run a query or an HTTP call |
| Anything answering "which rows count?" or "which value wins?" | Functions that only write a file or a table |
| Edge cases: no data, missing column, duplicate record | The SDK, pandas, the database driver |
- A suite that needs a database, a credential or a VPN will be skipped. If a test needs one, the function under test is doing too much — split it (rule 3) instead of mocking around it.
- Keep fixtures small. Three or four rows, only the columns the function reads. A fixture copied from a real export is one nobody dares to change and whose failures nobody can read.
- Tests also pin behaviour nobody is sure about. When a rule looks like it might be a bug but changing it would move published figures, write the test that states what it does today. A comment does not fail when someone "corrects" it; a test does.
Each table a pipeline publishes has a contract: its columns, the kind of values they hold, which may be empty, and what identifies a row. The contract is declared once and applied when the table is written.
Without one, the column list is retyped in every place that touches the table, the copies drift, and a renamed column reaches the dashboard as a broken visual rather than a failed run.
- Apply it at the boundary, immediately before publishing — not scattered through the transformations.
- Shape first, then content. Conforming a dataframe to the contract guarantees the agreed columns in the agreed order; validating it goes further and fails the run on a missing required value or a duplicated key. Start with the first, move to the second once you trust the data.
- A contract is shared, not per-pipeline. When two pipelines write the same table they must agree on its shape, which means reading the same declaration (rule 6).
- Test the contract itself, or applying it is a placebo: check that it really catches a missing column, a stray index column, an empty required field and a duplicated key.
Code that is not about one pipeline in particular — API clients, shared readers, schema contracts — belongs in its own repository, versioned and installed with pip. Copying it between pipelines guarantees the copies drift, and a fix then reaches one pipeline instead of all of them.
- Pin a version, not a branch. A pipeline must not change behaviour because someone pushed to
the library's
main. Pinning also makes a library upgrade a deliberate, reviewable change. - A library has its own tests and its own CI. If it is worth sharing, it is worth checking before it is shared.
- Keep the layering shallow inside a pipeline. One level down is easy to follow; a chain where each file re-exports the one below means reading three files to find where a function is defined.
- If a separate repository is too much right now, a package at the root of the repository, imported explicitly, already gets the ownership right — extracting it later is a file move.
Pipeline frameworks let you mark steps as tasks so they can be scheduled and run concurrently. Only do that when steps genuinely can run at the same time.
When each step consumes the result of the one before, marking them as tasks changes nothing about the order and costs clarity: the functions return futures instead of values, failures land one layer deeper in the traceback, and every reader has to work out whether anything is actually concurrent.
- Sequential steps are plain function calls.
- A task earns its place where there is real fan-out — the same work repeated over many independent inputs, with a single step at the end that needs all of them.
- Fan-out needs the body to be one function with explicit inputs, which is what rule 3 produces anyway.
Much of the code written from now on will be written with an AI agent. An agent that does not know
the pipeline cannot run locally will suggest running it; one that does not know a column name is
read by a dashboard will "fix" it. CLAUDE.md is the file both the agent and the new colleague
read first.
- Generate a first draft, then curate it.
/initin Claude Code proposes one. Delete anything you cannot verify, and anything the code already says. - Write down what the code cannot say: the domain vocabulary, where the pipeline runs, which names are an external interface, which odd-looking special cases encode a real incident.
- Give the exact commands — install, lint, format, test, run one test — copy-pasteable.
- State conventions as instructions, not values. "New logic goes in a function that takes data and returns data, with a test in the same change" is actionable; "we value testing" is not.
- List the traps, and say what is out of bounds: what not to touch, never commit a secret.
- Keep it to a screen or two, review it like code, and update it in the commit that invalidates it. It is context, not documentation; a manual nobody maintains is worse than nothing.