Skip to content

Commit 9c6c4aa

Browse files
Merge pull request #137 from FreshCode-Org/feature/validation-gauntlet-jwd
Validation Gauntlet: gold-labelled disposition benchmark + six defect fixes
2 parents 30d82b5 + 842a1d3 commit 9c6c4aa

20 files changed

Lines changed: 2139 additions & 24 deletions

‎.github/workflows/gauntlet.yml‎

Lines changed: 59 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,59 @@
1+
# Validation Gauntlet: gold-labelled disposition benchmark for the validation,
2+
# domain and text-cleaning surfaces. Runs the lightweight fixtures on every PR
3+
# and gates on the absolute thresholds plus no-regression vs the stored
4+
# baseline (benchmarks/gauntlet/baseline.json). Heavier sizes stay manual.
5+
name: Validation Gauntlet
6+
7+
on:
8+
pull_request:
9+
paths-ignore:
10+
- "docs/**"
11+
- "*.md"
12+
workflow_dispatch:
13+
inputs:
14+
rows:
15+
description: "rows per fixture"
16+
default: "300"
17+
update_baseline:
18+
description: "re-pin baseline.json from this run (commit it manually)"
19+
type: boolean
20+
default: false
21+
22+
permissions:
23+
contents: read
24+
25+
jobs:
26+
gauntlet:
27+
runs-on: ubuntu-latest
28+
timeout-minutes: 20
29+
steps:
30+
- uses: actions/checkout@v4
31+
- uses: actions/setup-python@v5
32+
with:
33+
python-version: "3.12"
34+
cache: pip
35+
- name: Install
36+
run: |
37+
python -m pip install --upgrade pip
38+
pip install -e ".[dev]"
39+
- name: Run gauntlet with gates
40+
run: |
41+
ROWS="${{ github.event.inputs.rows || '300' }}"
42+
EXTRA=""
43+
if [ "${{ github.event.inputs.update_baseline }}" = "true" ]; then
44+
EXTRA="--update-baseline"
45+
fi
46+
python -m benchmarks.gauntlet run --rows "$ROWS" --check $EXTRA
47+
- name: Job summary
48+
if: always()
49+
run: |
50+
if [ -f benchmarks/gauntlet/results/gauntlet.md ]; then
51+
cat benchmarks/gauntlet/results/gauntlet.md >> "$GITHUB_STEP_SUMMARY"
52+
fi
53+
- name: Upload results
54+
if: always()
55+
uses: actions/upload-artifact@v4
56+
with:
57+
name: gauntlet-results
58+
path: benchmarks/gauntlet/results/
59+
if-no-files-found: warn

‎.gitignore‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -34,3 +34,4 @@ crates/freshcore/target/
3434
# dev artifact (regenerated by running teacher tasks), not committed content.
3535
training/cache/*
3636
!training/cache/.gitkeep
37+
benchmarks/gauntlet/results/

‎CHANGELOG.md‎

Lines changed: 56 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,7 +6,63 @@ adheres to [Semantic Versioning](https://semver.org/).
66

77
## [Unreleased]
88

9+
### Added
10+
- **Validation Gauntlet** (`benchmarks/gauntlet/`, `docs/validation-gauntlet.md`):
11+
a gold-labelled disposition benchmark for the validation, domain and
12+
text-cleaning surfaces. Five deterministic fixtures (finance, healthcare,
13+
CRM, e-commerce, adversarial text) label every injected defect with the
14+
disposition FreshData should choose (preserve / repair / flag / review) and
15+
the harness scores detection P/R/F1, repair accuracy, review routing,
16+
preservation, corruption, escapes, false positives, audit completeness,
17+
determinism, trust monotonicity and runtime/memory. Runs on every PR
18+
(`gauntlet.yml`) with absolute gates plus no-regression checks against the
19+
stored `baseline.json`.
20+
- `CleanReport.coerced_cells`: per-cell record (`{column: {row: original}}`)
21+
of values that `fix_dtypes` nulled because they did not parse as the
22+
column's inferred type — the recovery source for quarantined cells, also
23+
included in `report.to_dict()`.
24+
- Date-field range validation in `fd.validate_fields`: `FieldSpec.min_value`
25+
/ `max_value` now accept a date string or timestamp for `date`/`datetime`
26+
fields, so a future date of birth or an 1875 admission date is flagged as a
27+
`domain_mismatch` (gauntlet finding).
28+
- Case-variant vocabulary suggestions in `fd.validate_fields`: a value that
29+
matches an `allowed_values` entry except for case (`ACTIVE` vs `active`) is
30+
no longer silently accepted — it gets a warning-severity issue with the
31+
canonical form as `suggestion` and action `accept_with_warning` (gauntlet
32+
finding).
33+
934
### Fixed
35+
- **Unparseable values are quarantined, never fabricated** (gauntlet finding,
36+
the `'apple'`-in-a-price-column case): when `fix_dtypes` converts a
37+
mostly-numeric (or datetime) text column, cells that fail to parse used to
38+
become `NaN` and then be silently imputed by the auto engine — turning
39+
junk into a fabricated median. They now stay missing, are excluded from
40+
auto-imputation, keep their originals in `report.coerced_cells`, and the
41+
decision is a `human_review` action in the audit trail. Genuine missing
42+
values (true `NaN`, sentinels like `"N/A"`) keep the documented
43+
auto-impute behaviour, and an explicit `impute=` request still fills
44+
everything.
45+
- Formatted-number stragglers (`"$1,234.56"`, `"1,200,500.00"`) in a
46+
mostly-plain numeric column are now parsed by the existing locale-aware
47+
rescue instead of being coerced to missing — the rescue previously only
48+
engaged when the plain parse failed the threshold entirely (gauntlet
49+
finding).
50+
- `fd.validate_fields` consensus inference now honours the same
51+
contamination boundary as the `fix_dtypes` warning that points users at it
52+
(dominant share ≥ 60% with at most a handful of stragglers). Previously
53+
the warning fired from a 60% parse share but the consensus gate required
54+
80%, so the exact frame the warning named sailed through
55+
`validate_fields` silently (gauntlet finding).
56+
- Explicitly allowed values are no longer swallowed by null-marker
57+
heuristics in `fd.validate_fields`: with
58+
`FieldSpec(allowed_values={"US", "DE", "NA"})`, `"NA"` is Namibia, not a
59+
missing value (gauntlet finding).
60+
- `clean_text` / `validate_fields` text normalization no longer rewrites
61+
typography in content-bearing fields: for `free_text`, `text` and entity
62+
name types, the punctuation→ASCII mapping (curly quotes, em-dashes, prime
63+
marks — `12″` became `12"`) is withheld, matching the field-aware safety
64+
contract. Untyped columns keep the existing behaviour (gauntlet finding).
65+
1066
- `anonymize()` called with no `rules` and no `detection_config` now emits
1167
a `UserWarning` instead of silently returning the data unchanged — a
1268
privacy call that does nothing must say so. Behavior is otherwise

‎benchmarks/README.md‎

Lines changed: 6 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -9,6 +9,12 @@ pyjanitor baselines.
99
> The harness calls FreshData exactly as a user would. It never modifies library
1010
> internals.
1111
12+
> Sibling harness: `benchmarks/gauntlet/` (the **Validation Gauntlet**)
13+
> scores per-cell dispositions — preserve / repair / flag / review — for
14+
> the validation, domain and text-cleaning surfaces against gold labels,
15+
> and gates every PR via `.github/workflows/gauntlet.yml`. See
16+
> `docs/validation-gauntlet.md`.
17+
1218
## Layout
1319

1420
```

‎benchmarks/gauntlet/__init__.py‎

Lines changed: 28 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,28 @@
1+
"""FreshData Validation Gauntlet.
2+
3+
Gold-labelled adversarial fixtures plus a harness that measures how
4+
FreshData's validation surfaces (``fd.clean``, ``fd.validate_fields``,
5+
``fd.clean_text``, domain packs, the semantic layer, PII detection) treat
6+
each labelled cell: preserve, repair, flag, or route to review.
7+
8+
Unlike CleanBench (which scores whole-frame repair fidelity against a clean
9+
oracle), the gauntlet scores *dispositions*: every injected defect carries the
10+
disposition FreshData should choose, and every adversarial trap is a valid
11+
value that must survive cleaning untouched.
12+
13+
Run ``python -m benchmarks.gauntlet run`` from the repo root.
14+
"""
15+
16+
from .fixtures import FIXTURES, GauntletFixture, GoldCell, build_fixture
17+
from .metrics import compute_metrics
18+
from .runner import run_fixture, run_gauntlet
19+
20+
__all__ = [
21+
"FIXTURES",
22+
"GauntletFixture",
23+
"GoldCell",
24+
"build_fixture",
25+
"compute_metrics",
26+
"run_fixture",
27+
"run_gauntlet",
28+
]

‎benchmarks/gauntlet/__main__.py‎

Lines changed: 67 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,67 @@
1+
"""Validation Gauntlet CLI.
2+
3+
Run from the repo root::
4+
5+
python -m benchmarks.gauntlet run # run + write results/
6+
python -m benchmarks.gauntlet run --check # also gate (CI mode)
7+
python -m benchmarks.gauntlet run --update-baseline
8+
"""
9+
10+
from __future__ import annotations
11+
12+
import argparse
13+
import json
14+
import sys
15+
from pathlib import Path
16+
17+
from .fixtures import DEFAULT_ROWS, DEFAULT_SEED, FIXTURES
18+
from .metrics import compute_metrics
19+
from .report import check_gates, render_markdown, results_payload, write_json
20+
from .runner import run_gauntlet
21+
22+
RESULTS_DIR = Path(__file__).parent / "results"
23+
BASELINE_PATH = Path(__file__).parent / "baseline.json"
24+
25+
26+
def main(argv: list[str] | None = None) -> int:
27+
parser = argparse.ArgumentParser(prog="python -m benchmarks.gauntlet")
28+
sub = parser.add_subparsers(dest="command", required=True)
29+
run_p = sub.add_parser("run", help="run the gauntlet and write JSON + Markdown")
30+
run_p.add_argument("--rows", type=int, default=DEFAULT_ROWS)
31+
run_p.add_argument("--seed", type=int, default=DEFAULT_SEED)
32+
run_p.add_argument("--fixtures", nargs="*", choices=sorted(FIXTURES))
33+
run_p.add_argument("--check", action="store_true",
34+
help="exit 1 when a gate fails or the baseline regresses")
35+
run_p.add_argument("--update-baseline", action="store_true",
36+
help="write this run as the stored baseline")
37+
args = parser.parse_args(argv)
38+
39+
runs = run_gauntlet(n_rows=args.rows, seed=args.seed, fixtures=args.fixtures)
40+
metrics = {name: compute_metrics(r) for name, r in runs.items()}
41+
payload = results_payload(metrics, n_rows=args.rows, seed=args.seed)
42+
43+
write_json(payload, RESULTS_DIR / "gauntlet.json")
44+
markdown = render_markdown(payload)
45+
(RESULTS_DIR / "gauntlet.md").write_text(markdown)
46+
print(markdown)
47+
print(f"results: {RESULTS_DIR / 'gauntlet.json'}")
48+
49+
if args.update_baseline:
50+
write_json(payload, BASELINE_PATH)
51+
print(f"baseline updated: {BASELINE_PATH}")
52+
53+
if args.check:
54+
baseline = (json.loads(BASELINE_PATH.read_text())
55+
if BASELINE_PATH.exists() else None)
56+
problems = check_gates(payload, baseline)
57+
if problems:
58+
print("\nGATE FAILURES:", file=sys.stderr)
59+
for p in problems:
60+
print(f" - {p}", file=sys.stderr)
61+
return 1
62+
print("all gates passed")
63+
return 0
64+
65+
66+
if __name__ == "__main__":
67+
raise SystemExit(main())

0 commit comments

Comments
 (0)