Skip to content

Commit a21e19c

Browse files
Merge pull request #142 from FreshCode-Org/feature/truthbench-jwd
feat: add TruthBench semantic red-team foundation
2 parents a1da862 + 0eae829 commit a21e19c

43 files changed

Lines changed: 10430 additions & 0 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.superpowers/sdd/task-4-report.md‎

Lines changed: 110 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,110 @@
1+
# Task 4 Report: Add finance, healthcare, retail, and CRM gold datasets
2+
3+
## Status
4+
5+
DONE. The four deterministic 16-row domain builders are registered and covered by focused and adjacent TruthBench tests.
6+
7+
## TDD evidence
8+
9+
### RED
10+
11+
```text
12+
PYTHONPATH=src python -m pytest tests/truthbench/test_fixtures.py -q --no-cov
13+
```
14+
15+
Observed 16 failures before implementation: each new domain was absent from the fixture registry (`FixtureError: unknown fixture domain`).
16+
17+
### GREEN
18+
19+
```text
20+
PYTHONPATH=src python -m pytest tests/truthbench/test_fixtures.py -q --no-cov
21+
```
22+
23+
Result: `28 passed`.
24+
25+
```text
26+
PYTHONPATH=src python -m pytest tests/truthbench/test_fixtures.py tests/truthbench/test_models_exact.py tests/truthbench/test_schema.py -q --no-cov
27+
```
28+
29+
Result: all focused, model, and schema tests green.
30+
31+
Static checks:
32+
33+
```text
34+
python -m ruff check benchmarks/truthbench/fixtures tests/truthbench/test_fixtures.py
35+
python -m ruff format --check benchmarks/truthbench/fixtures tests/truthbench/test_fixtures.py
36+
mypy --ignore-missing-imports --explicit-package-bases benchmarks/truthbench/fixtures tests/truthbench/conftest.py tests/truthbench/test_fixtures.py
37+
git diff --check
38+
```
39+
40+
All passed.
41+
42+
## Files
43+
44+
- `benchmarks/truthbench/fixtures/finance.py` — finance frame and oracle labels for Apple semantic traps, numeric/currency/date defects, protected ticker conflict, and synthetic PII canaries.
45+
- `benchmarks/truthbench/fixtures/healthcare.py` — healthcare frame and oracle labels for ICD/LOINC, temperature and dose units, FHIR/partial/impossible dates, Unicode, and PHI canaries.
46+
- `benchmarks/truthbench/fixtures/retail.py` — retail frame and oracle labels for leading-zero identifiers, free/return quantities, locale currency formats, mojibake/HTML, multilingual values, and review PII.
47+
- `benchmarks/truthbench/fixtures/crm.py` — CRM frame and oracle labels for Unicode/contact ambiguity, lifecycle contradiction, zero-width/hidden PII, Apple semantic traps, and protected IDs.
48+
- `benchmarks/truthbench/fixtures/__init__.py` — registry entries and stable domain order.
49+
- `tests/truthbench/test_fixtures.py` — domain completeness, disposition, deterministic-seed, and adversarial marker tests.
50+
51+
## Self-review
52+
53+
- Every builder emits exactly 16 rows with stable string indexes, fixed `2026-01-15` UTC reference metadata, explicit locale metadata, complete physical-cell labels, and at least 12 injected adversarial cells.
54+
- All four dispositions are represented per domain. Row cases cover exact duplicates and removed rows; schema cases cover added, removed, renamed, reordered, and type-drifted columns without mislabeling absent cells.
55+
- Sensitive values are whole-value synthetic `.invalid`, `TB-*`, or `555-01xx` forms, preserving the base builder's redaction and canary invariants. Zero-width examples are intentionally non-sensitive representation traps.
56+
- Seed values are included only in a deterministic batch marker; same-seed builds produce byte-stable serialized oracle payloads and identical fixture hashes for seeds `1729` and `2718`.
57+
- No FreshData runtime, LLM/provider, network, or external data dependency is used.
58+
59+
## Concerns
60+
61+
- The legacy `minimal` fixture remains registered for backwards compatibility; the four Task 4 domains are appended in stable order. A later registry task may choose to retire `minimal` once its callers migrate.
62+
- Row/schema expectations are metadata cases (the physical frame remains rectangular), consistent with the oracle contract that removed rows/columns cannot have cell labels.
63+
64+
## Review follow-up: healthcare reference codes and content assertions
65+
66+
### RED
67+
68+
After review, focused content tests were strengthened to inspect each actual `GoldCell.family`, disposition, and adversarial frame value. The first run exposed four failures: the healthcare rare-code test rejected the placeholder `G rare`; finance, retail, and CRM injection-count assertions correctly counted only non-preserve dispositions rather than all injected cells.
69+
70+
```text
71+
PYTHONPATH=src python -m pytest tests/truthbench/test_fixtures.py -q --no-cov
72+
```
73+
74+
Result before fixes: `4 failed, 24 passed`.
75+
76+
### GREEN
77+
78+
Healthcare preserve cases now use values present in the bundled reference sets (`Z79.4`, `F17.210`, and `9843-4`), and tests load those references plus validate ICD/LOINC syntax. Content tests assert the finance, healthcare, retail, and CRM families directly against physical cells; injected-cell counts use non-background families, including preserve injections.
79+
80+
```text
81+
PYTHONPATH=src python -m pytest tests/truthbench/test_fixtures.py -q --no-cov
82+
```
83+
84+
Result: `28 passed`.
85+
86+
Complete adjacent verification (`tests/truthbench/test_fixtures.py`, `tests/truthbench/test_models_exact.py`, and `tests/truthbench/test_schema.py`) passed. Ruff check/format, mypy, and `git diff --check` also passed.
87+
88+
## Review follow-up: contract-gap family coverage
89+
90+
### RED
91+
92+
Added direct assertions for the finance USD/EUR/INR conflict and zero-width memo, healthcare protected-DOB repair conflict and MRN tail canary, and exact row/schema family sets for finance, healthcare, and CRM. The initial run exposed the missing `zero-width-memo` family label (the value existed but was grouped under the broader invisible-PII family).
93+
94+
```text
95+
PYTHONPATH=src python -m pytest tests/truthbench/test_fixtures.py -q --no-cov
96+
```
97+
98+
Result: `1 failed, 30 passed` (`StopIteration` while locating the required zero-width memo family).
99+
100+
### GREEN
101+
102+
The zero-width memo cell now has its own `zero-width-memo` family; all assertions inspect actual frame values, dispositions, sensitivity, and exact case-family sets.
103+
104+
```text
105+
PYTHONPATH=src python -m pytest tests/truthbench/test_fixtures.py -q --no-cov
106+
```
107+
108+
Result: `31 passed`.
109+
110+
Complete fixture/models/schema verification passed (`... passed`), as did Ruff check/format, mypy, and `git diff --check`.

‎.superpowers/sdd/task-5-report.md‎

Lines changed: 85 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,85 @@
1+
# Task 5 Report: Complete eight-domain TruthBench corpus
2+
3+
## Status
4+
5+
DONE. Logistics, government, education, and insurance now have deterministic 16-row gold fixtures. The registry is the stable alphabetical eight-domain order (`crm`, `education`, `finance`, `government`, `healthcare`, `insurance`, `logistics`, `retail`) with the temporary minimal registry removed.
6+
7+
## TDD evidence
8+
9+
### RED
10+
11+
After adding domain-content tests, the focused fixture suite failed with the expected missing-builder errors for all four new domains (`FixtureError: unknown fixture domain`). The failure run was:
12+
13+
```text
14+
PYTHONPATH=src python -m pytest tests/truthbench/test_fixtures.py -q --no-cov
15+
```
16+
17+
### GREEN
18+
19+
The focused suite now passes:
20+
21+
```text
22+
PYTHONPATH=src python -m pytest tests/truthbench/test_fixtures.py -q --no-cov
23+
```
24+
25+
Result: `47 passed`.
26+
27+
Adjacent fixture/model/schema verification:
28+
29+
```text
30+
PYTHONPATH=src python -m pytest \
31+
tests/truthbench/test_fixtures.py \
32+
tests/truthbench/test_models_exact.py \
33+
tests/truthbench/test_schema.py -q --no-cov
34+
```
35+
36+
Result: `112 passed`.
37+
38+
Static checks:
39+
40+
```text
41+
python -m ruff check benchmarks/truthbench/fixtures tests/truthbench/test_fixtures.py
42+
python -m ruff format --check benchmarks/truthbench/fixtures tests/truthbench/test_fixtures.py
43+
mypy --ignore-missing-imports --explicit-package-bases \
44+
benchmarks/truthbench/fixtures tests/truthbench/test_fixtures.py
45+
git diff --check
46+
```
47+
48+
All passed.
49+
50+
## Coverage
51+
52+
- `logistics.py` covers valid UN/LOCODE-like references, kg/lb and C/F units, cross-timezone windows, 24:00 transport values, address PII, late tracking, and protected shipment IDs.
53+
- `government.py` covers leading-zero IDs, Indian/international grouping, fiscal/calendar ambiguity, multilingual labels, restricted national IDs, mixed legacy encoding, and retention/repair policy contradiction.
54+
- `education.py` covers student IDs, letter/percentage/GPA scales, school-year ambiguity, zero scores, enrollment ordering, guardian contacts, FERPA notes, and protected grade-policy conflict.
55+
- `insurance.py` covers policy/claim IDs, premium/reserve currency mismatch, negative reserve review, incident/report ordering, state contradiction, claimant/medical PII, and protected policy numbers.
56+
- All four builders emit complete physical-cell labels, all four dispositions, row duplicate/removal cases, five schema drift cases, fixed UTC/reference metadata, deterministic seed batches, and privacy-safe synthetic canaries.
57+
- Corpus-level tests assert all required trap categories occur across the eight domains.
58+
59+
## Concerns
60+
61+
None. No FreshData runtime, LLM/provider, network, or external data dependency is used.
62+
63+
## Review follow-up: explicit contract coverage
64+
65+
### RED
66+
67+
The new required-family contract test intentionally used the repaired numeric output (`95`) as the adversarial frame value for education `edu-07`. The focused test failed because the actual frame value is the required raw value `"95%"` while the repair oracle separately stores `95.0` as its expected output.
68+
69+
```text
70+
PYTHONPATH=src python -m pytest \
71+
tests/truthbench/test_fixtures.py::test_required_domain_families_match_actual_values_and_dispositions \
72+
-q --no-cov
73+
```
74+
75+
Observed: one failure at `edu-07` (`'95%' != 95`).
76+
77+
### GREEN
78+
79+
The contract now asserts the raw value, family, disposition, and typed repair output. It also uses an explicit per-domain family mapping covering every required category (including logistics lb/F units, government IDs/grouping/language/fiscal/protected case, education scales/contacts/protected grade, and insurance IDs/grouped premium/PII/protected policy).
80+
81+
```text
82+
PYTHONPATH=src python -m pytest tests/truthbench/test_fixtures.py -q --no-cov
83+
```
84+
85+
Result: `48 passed`.

‎.superpowers/sdd/task-6-report.md‎

Lines changed: 202 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,202 @@
1+
# Task 6 Report: Privacy-safe values and exhaustive sink scanning
2+
3+
## Status
4+
5+
DONE. `SinkScanner` now scans normalized canary variants across nested TruthBench
6+
sinks and emits only `Leak(canary_id, variant, path)` metadata. Redaction markers
7+
contain run-scoped HMAC-SHA256 digests and never matched text.
8+
9+
## TDD evidence
10+
11+
### RED
12+
13+
```text
14+
PYTHONPATH=src python -m pytest tests/truthbench/test_privacy.py -q --no-cov
15+
```
16+
17+
Observed collection failure before implementation:
18+
19+
```text
20+
ModuleNotFoundError: No module named 'benchmarks.truthbench.privacy'
21+
```
22+
23+
### GREEN
24+
25+
Focused privacy suite:
26+
27+
```text
28+
PYTHONPATH=src python -m pytest tests/truthbench/test_privacy.py -q --no-cov
29+
```
30+
31+
Result: `18 passed`.
32+
33+
Adjacent TruthBench suites:
34+
35+
```text
36+
PYTHONPATH=src python -m pytest \
37+
tests/truthbench/test_privacy.py \
38+
tests/truthbench/test_fixtures.py \
39+
tests/truthbench/test_models_exact.py \
40+
tests/truthbench/test_schema.py -q --no-cov
41+
```
42+
43+
Result: `131 passed`.
44+
45+
Static checks:
46+
47+
```text
48+
python -m ruff check benchmarks/truthbench/privacy.py \
49+
benchmarks/truthbench/__init__.py tests/truthbench/test_privacy.py
50+
python -m ruff format --check benchmarks/truthbench/privacy.py \
51+
benchmarks/truthbench/__init__.py tests/truthbench/test_privacy.py
52+
mypy --ignore-missing-imports --explicit-package-bases \
53+
benchmarks/truthbench/privacy.py tests/truthbench/test_privacy.py
54+
git diff --check
55+
```
56+
57+
## Review follow-up: MultiIndex and hostile-label privacy hardening
58+
59+
### RED
60+
61+
Added regression tests for MultiIndex columns/index level names, tuple-label
62+
structure during redaction, and custom non-string labels whose stringification
63+
contains a canary. Before the fix:
64+
65+
```text
66+
pytest -q tests/truthbench/test_privacy.py --no-cov
67+
```
68+
69+
Result: `24 passed, 2 failed`.
70+
71+
The failures were the expected defects: redacting tuple labels converted them
72+
to lists and raised pandas `ValueError` (length mismatch), while hostile custom
73+
labels were not scanned.
74+
75+
### GREEN
76+
77+
Structure-preserving label traversal/redaction now handles tuple/MultiIndex
78+
levels and names, and sanitizes custom label paths before reporting or replacing
79+
them with digest markers:
80+
81+
```text
82+
pytest -q tests/truthbench/test_privacy.py --no-cov
83+
```
84+
85+
Result: `26 passed`.
86+
87+
Adjacent TruthBench verification:
88+
89+
```text
90+
pytest -q tests/truthbench --no-cov
91+
```
92+
93+
Result: `139 passed`.
94+
95+
Static checks:
96+
97+
```text
98+
ruff check benchmarks/truthbench/privacy.py tests/truthbench/test_privacy.py
99+
ruff format --check benchmarks/truthbench/privacy.py tests/truthbench/test_privacy.py
100+
mypy benchmarks/truthbench/privacy.py
101+
git diff --check
102+
```
103+
104+
All passed.
105+
106+
All passed.
107+
108+
## Files
109+
110+
- `benchmarks/truthbench/privacy.py` — `Leak`, `PrivacySafeValue`, named
111+
normalizers, run-scoped HMAC scanner, recursive redaction, typed redaction
112+
handling, pandas/dataclass/bytes support, self-test, and named sink entry points.
113+
- `benchmarks/truthbench/__init__.py` — exports privacy scanner primitives.
114+
- `tests/truthbench/test_privacy.py` — mutation matrix for every required
115+
normalized form, nested sink coverage, redaction/self-test behavior, and typed
116+
redaction digest safety.
117+
118+
## Self-review
119+
120+
- Leak objects contain only identifiers, transform labels, and JSONPath-like
121+
locations; `repr(leaks)` cannot repeat canary text.
122+
- Literal, case-folded, whitespace-stripped, punctuation-stripped, digit-only,
123+
URL-decoded, HTML-unescaped, UTF-8/hex/escape bytes, NFKC/NFC/NFD,
124+
zero-width-removed, and JSON-escaped forms are covered.
125+
- Mapping, sequence, dataclass, pandas DataFrame/Series/Index, exception,
126+
report, plan, generated-code, stream, markup, JSON, and failure-artifact sinks
127+
are traversed. Exact redacted `TypedValue` payloads are treated as safe and
128+
their HMAC digests are not re-scanned as plaintext.
129+
- Redaction is recursive, emits `[REDACTED:<digest>]`, and `self_test` raises on
130+
an unredacted result while accepting the scanner's own redacted output.
131+
132+
## Concerns
133+
134+
- The scanner intentionally treats digit-only normalized forms conservatively;
135+
very short numeric canaries can match unrelated text. Fixtures use synthetic
136+
identifiers and email/phone canaries, so this does not affect the bundled
137+
corpus.
138+
- Redacting a pandas object may change a sensitive column to object/string dtype,
139+
which is preferable to retaining a raw value in an audit sink.
140+
141+
## Follow-up RED/GREEN evidence
142+
143+
A follow-up regression test used the exact sensitive `TypedValue` redaction shape
144+
with a one-digit canary. Before the redacted-payload guard, the digest's hex text
145+
was incorrectly reported as a digit-only leak. After adding the guard:
146+
147+
```text
148+
PYTHONPATH=src python -m pytest \
149+
tests/truthbench/test_privacy.py::test_scanner_accepts_exact_typed_redaction_without_scanning_digest \
150+
-q --no-cov
151+
```
152+
153+
Result: `1 passed`.
154+
155+
## Review follow-up: marker, key, and label hardening
156+
157+
### RED
158+
159+
The review regression matrix was run before the hardening changes:
160+
161+
```text
162+
PYTHONPATH=src python -m pytest tests/truthbench/test_privacy.py -q --no-cov
163+
```
164+
165+
Result: `4 failed, 17 passed` for forged redaction markers, sensitive mapping
166+
keys, pandas labels, and the digest compatibility alias.
167+
168+
### GREEN
169+
170+
After validating markers against the scanner's own 64-hex HMAC digest set,
171+
redacting mapping keys and pandas column/index/name labels, scanning arbitrary
172+
key stringifications, and adding `digest` as an alias:
173+
174+
```text
175+
PYTHONPATH=src python -m pytest tests/truthbench/test_privacy.py -q --no-cov
176+
```
177+
178+
Result: `23 passed`.
179+
180+
Adjacent TruthBench verification:
181+
182+
```text
183+
PYTHONPATH=src python -m pytest \
184+
tests/truthbench/test_privacy.py \
185+
tests/truthbench/test_fixtures.py \
186+
tests/truthbench/test_models_exact.py \
187+
tests/truthbench/test_schema.py -q --no-cov
188+
```
189+
190+
Result: `136 passed`.
191+
192+
Static checks were rerun after the follow-up and remained clean:
193+
194+
```text
195+
python -m ruff check benchmarks/truthbench/privacy.py \
196+
tests/truthbench/test_privacy.py
197+
python -m ruff format --check benchmarks/truthbench/privacy.py \
198+
tests/truthbench/test_privacy.py
199+
mypy --ignore-missing-imports --explicit-package-bases \
200+
benchmarks/truthbench/privacy.py tests/truthbench/test_privacy.py
201+
git diff --check
202+
```

0 commit comments

Comments
 (0)