Skip to content

Commit 842a1d3

Browse files
ci+docs: run the Validation Gauntlet on every PR; document it
- gauntlet.yml: PR-triggered lightweight run (300 rows) with absolute gates and no-regression checks vs the committed baseline.json, Markdown job summary, uploaded artifacts; heavier sizes via workflow_dispatch. - docs/validation-gauntlet.md (+ mkdocs nav, benchmarks/README pointer). - CHANGELOG: gauntlet, coerced-cell quarantine, formatted-number rescue, fieldcheck consensus/vocabulary/case/date-bounds fixes, typography protection. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
1 parent d92bc4f commit 842a1d3

5 files changed

Lines changed: 195 additions & 0 deletions

File tree

‎.github/workflows/gauntlet.yml‎

Lines changed: 59 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,59 @@
1+
# Validation Gauntlet: gold-labelled disposition benchmark for the validation,
2+
# domain and text-cleaning surfaces. Runs the lightweight fixtures on every PR
3+
# and gates on the absolute thresholds plus no-regression vs the stored
4+
# baseline (benchmarks/gauntlet/baseline.json). Heavier sizes stay manual.
5+
name: Validation Gauntlet
6+
7+
on:
8+
pull_request:
9+
paths-ignore:
10+
- "docs/**"
11+
- "*.md"
12+
workflow_dispatch:
13+
inputs:
14+
rows:
15+
description: "rows per fixture"
16+
default: "300"
17+
update_baseline:
18+
description: "re-pin baseline.json from this run (commit it manually)"
19+
type: boolean
20+
default: false
21+
22+
permissions:
23+
contents: read
24+
25+
jobs:
26+
gauntlet:
27+
runs-on: ubuntu-latest
28+
timeout-minutes: 20
29+
steps:
30+
- uses: actions/checkout@v4
31+
- uses: actions/setup-python@v5
32+
with:
33+
python-version: "3.12"
34+
cache: pip
35+
- name: Install
36+
run: |
37+
python -m pip install --upgrade pip
38+
pip install -e ".[dev]"
39+
- name: Run gauntlet with gates
40+
run: |
41+
ROWS="${{ github.event.inputs.rows || '300' }}"
42+
EXTRA=""
43+
if [ "${{ github.event.inputs.update_baseline }}" = "true" ]; then
44+
EXTRA="--update-baseline"
45+
fi
46+
python -m benchmarks.gauntlet run --rows "$ROWS" --check $EXTRA
47+
- name: Job summary
48+
if: always()
49+
run: |
50+
if [ -f benchmarks/gauntlet/results/gauntlet.md ]; then
51+
cat benchmarks/gauntlet/results/gauntlet.md >> "$GITHUB_STEP_SUMMARY"
52+
fi
53+
- name: Upload results
54+
if: always()
55+
uses: actions/upload-artifact@v4
56+
with:
57+
name: gauntlet-results
58+
path: benchmarks/gauntlet/results/
59+
if-no-files-found: warn

‎CHANGELOG.md‎

Lines changed: 56 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,63 @@ adheres to [Semantic Versioning](https://semver.org/).
66

77
## [Unreleased]
88

9+
### Added
10+
- **Validation Gauntlet** (`benchmarks/gauntlet/`, `docs/validation-gauntlet.md`):
11+
a gold-labelled disposition benchmark for the validation, domain and
12+
text-cleaning surfaces. Five deterministic fixtures (finance, healthcare,
13+
CRM, e-commerce, adversarial text) label every injected defect with the
14+
disposition FreshData should choose (preserve / repair / flag / review) and
15+
the harness scores detection P/R/F1, repair accuracy, review routing,
16+
preservation, corruption, escapes, false positives, audit completeness,
17+
determinism, trust monotonicity and runtime/memory. Runs on every PR
18+
(`gauntlet.yml`) with absolute gates plus no-regression checks against the
19+
stored `baseline.json`.
20+
- `CleanReport.coerced_cells`: per-cell record (`{column: {row: original}}`)
21+
of values that `fix_dtypes` nulled because they did not parse as the
22+
column's inferred type — the recovery source for quarantined cells, also
23+
included in `report.to_dict()`.
24+
- Date-field range validation in `fd.validate_fields`: `FieldSpec.min_value`
25+
/ `max_value` now accept a date string or timestamp for `date`/`datetime`
26+
fields, so a future date of birth or an 1875 admission date is flagged as a
27+
`domain_mismatch` (gauntlet finding).
28+
- Case-variant vocabulary suggestions in `fd.validate_fields`: a value that
29+
matches an `allowed_values` entry except for case (`ACTIVE` vs `active`) is
30+
no longer silently accepted — it gets a warning-severity issue with the
31+
canonical form as `suggestion` and action `accept_with_warning` (gauntlet
32+
finding).
33+
934
### Fixed
35+
- **Unparseable values are quarantined, never fabricated** (gauntlet finding,
36+
the `'apple'`-in-a-price-column case): when `fix_dtypes` converts a
37+
mostly-numeric (or datetime) text column, cells that fail to parse used to
38+
become `NaN` and then be silently imputed by the auto engine — turning
39+
junk into a fabricated median. They now stay missing, are excluded from
40+
auto-imputation, keep their originals in `report.coerced_cells`, and the
41+
decision is a `human_review` action in the audit trail. Genuine missing
42+
values (true `NaN`, sentinels like `"N/A"`) keep the documented
43+
auto-impute behaviour, and an explicit `impute=` request still fills
44+
everything.
45+
- Formatted-number stragglers (`"$1,234.56"`, `"1,200,500.00"`) in a
46+
mostly-plain numeric column are now parsed by the existing locale-aware
47+
rescue instead of being coerced to missing — the rescue previously only
48+
engaged when the plain parse failed the threshold entirely (gauntlet
49+
finding).
50+
- `fd.validate_fields` consensus inference now honours the same
51+
contamination boundary as the `fix_dtypes` warning that points users at it
52+
(dominant share ≥ 60% with at most a handful of stragglers). Previously
53+
the warning fired from a 60% parse share but the consensus gate required
54+
80%, so the exact frame the warning named sailed through
55+
`validate_fields` silently (gauntlet finding).
56+
- Explicitly allowed values are no longer swallowed by null-marker
57+
heuristics in `fd.validate_fields`: with
58+
`FieldSpec(allowed_values={"US", "DE", "NA"})`, `"NA"` is Namibia, not a
59+
missing value (gauntlet finding).
60+
- `clean_text` / `validate_fields` text normalization no longer rewrites
61+
typography in content-bearing fields: for `free_text`, `text` and entity
62+
name types, the punctuation→ASCII mapping (curly quotes, em-dashes, prime
63+
marks — `12″` became `12"`) is withheld, matching the field-aware safety
64+
contract. Untyped columns keep the existing behaviour (gauntlet finding).
65+
1066
- `anonymize()` called with no `rules` and no `detection_config` now emits
1167
a `UserWarning` instead of silently returning the data unchanged — a
1268
privacy call that does nothing must say so. Behavior is otherwise

‎benchmarks/README.md‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -9,6 +9,12 @@ pyjanitor baselines.
99
> The harness calls FreshData exactly as a user would. It never modifies library
1010
> internals.
1111
12+
> Sibling harness: `benchmarks/gauntlet/` (the **Validation Gauntlet**)
13+
> scores per-cell dispositions — preserve / repair / flag / review — for
14+
> the validation, domain and text-cleaning surfaces against gold labels,
15+
> and gates every PR via `.github/workflows/gauntlet.yml`. See
16+
> `docs/validation-gauntlet.md`.
17+
1218
## Layout
1319

1420
```

‎docs/validation-gauntlet.md‎

Lines changed: 73 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,73 @@
1+
# Validation Gauntlet
2+
3+
The Validation Gauntlet is a gold-labelled disposition benchmark for
4+
FreshData's validation surfaces: `fd.clean`, `fd.validate_fields`,
5+
`fd.clean_text`, the semantic layer, the domain packs, text-encoding linting
6+
and PII detection. It lives in `benchmarks/gauntlet/` and runs on every pull
7+
request (`.github/workflows/gauntlet.yml`).
8+
9+
Where [CleanBench](benchmarks.md) scores whole-frame repair fidelity against a
10+
clean oracle, the gauntlet scores **decisions**. Every injected problem cell
11+
carries the disposition FreshData should choose:
12+
13+
| disposition | meaning |
14+
|---|---|
15+
| `preserve` | valid (often unusual) data — must survive byte-identical, no error-severity issue |
16+
| `repair` | a safe deterministic repair exists — the gold value is known |
17+
| `flag` | must be detected, never auto-changed |
18+
| `review` | ambiguous — must be routed to quarantine / manual review, never guessed |
19+
20+
Automatic removal is never the default correct answer: a `flag` or `review`
21+
cell that the pipeline mutates counts as a **corruption**, the most severe
22+
verdict in the report.
23+
24+
## Fixtures
25+
26+
Five deterministic fixtures (seeded, 300 rows each by default) in
27+
`benchmarks/gauntlet/fixtures.py`: `finance`, `healthcare`, `crm`,
28+
`ecommerce` and `text` (adversarial free text). They cover missing values,
29+
impossible ranges, malformed and impossible dates, invalid identifiers,
30+
duplicates, numbers stored as text, currency/percent formats, unit confusion,
31+
casing and whitespace noise, misspellings, Unicode/encoding noise, emojis,
32+
HTML fragments, mixed-language values, PII, and adversarial traps designed to
33+
trigger false corrections (`X Æ A-12` as a name, `007` as an id, `NA` as a
34+
country, `None` as a brand token, `AB-` as a blood type, `BRK.B` as a ticker).
35+
36+
The flagship case: the string `apple` in a price column must be quarantined
37+
with its original value preserved in `report.coerced_cells` — never silently
38+
imputed — while `Apple` (company), `AAPL` (ticker) survive untouched and the
39+
lowercase ticker `apple` is routed to review with the suggestion `AAPL`.
40+
41+
## Metrics and gates
42+
43+
Per fixture: detection precision / recall / F1, repair accuracy (with the
44+
surface that produced each repair), review-routing rate, preservation rate,
45+
corruption count, escape rate, false-positive rate, audit completeness,
46+
determinism, trust-score monotonicity, wall-clock and peak memory.
47+
48+
CI fails a pull request when any fixture breaks the absolute gates
49+
(`benchmarks/gauntlet/report.py::GATES` — zero corruption, 100% preservation
50+
and audit completeness, F1 ≥ 0.85, repair accuracy ≥ 0.95, escapes ≤ 10%) or
51+
when detection recall/F1 drops below the stored baseline
52+
(`benchmarks/gauntlet/baseline.json`).
53+
54+
## Running locally
55+
56+
```bash
57+
python -m benchmarks.gauntlet run # JSON + Markdown into benchmarks/gauntlet/results/
58+
python -m benchmarks.gauntlet run --check # CI mode: exit 1 on gate failure
59+
python -m benchmarks.gauntlet run --rows 2000 # heavier run, manual only
60+
python -m benchmarks.gauntlet run --update-baseline
61+
```
62+
63+
Opt-in behaviour is graded separately from defaults: the safe `clean_text`
64+
config is never scored on lossy operations; HTML stripping and NFKC folding
65+
earn repair credit only from the explicit opt-in pass, and imputation credit
66+
for sentinel cells requires the fill to be audited.
67+
68+
## Known accepted gap
69+
70+
A hostile SQL-injection payload inside a `person_name` field escapes
71+
detection. Any plausibility heuristic tight enough to catch it false-positives
72+
on legally real names (`X Æ A-12`), so the gauntlet documents the escape
73+
rather than forcing a lossy check.

‎mkdocs.yml‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -136,6 +136,7 @@ nav:
136136
- Threat model: threat-model.md
137137
- Production-readiness checklist: production-readiness.md
138138
- Benchmarks: benchmarks.md
139+
- Validation Gauntlet: validation-gauntlet.md
139140
- Benchmark fixtures: fixtures.md
140141
- API Reference: api-reference.md
141142
- FAQ: faq.md

0 commit comments

Comments
 (0)