fix: report unmeasured quality-debt dimensions as not assessed; list columns aren't mixed types - #420
Merged
kevincostner17 merged 2 commits intoSep 15, 2026
Conversation
evaluate_quality_debt scored a dimension a clean 0.0 when it was never measured: type_instability when profiling failed, pii_risk when the PII scan was unavailable or failed, schema_drift and category_churn when no baseline was passed, and any dimension missing from the scores. Each item counted as assessed, so the gate could report pass on evidence it never collected, and the ledger stored a 0.0 that reset escalation history. These dimensions now use the DebtItem.assessed mechanism from #414: score and over_threshold serialise as None, the detail says why, the item is listed in gate.unassessed and shown as "? ... not assessed", and it adds nothing to the total, the gate status or the ledger. Runs with a baseline and a working profile and PII scan are byte-identical.
compute_trust_score tagged an object column holding Python lists or dicts as "mixed types" and lowered consistency (66.7 for one such column out of three), because infer_dtype returns "mixed" for any object column of containers. The equivalent nested Arrow list column infers as "unknown-array" and scored 100. A column whose non-null values are all lists, all dicts or all tuples is now uniform and scores like the Arrow equivalent. Columns that mix kinds (strings with numbers, lists with scalars, lists with dicts or tuples) are still flagged. The container check only runs when infer_dtype already reports mixed, so scores for plain frames are unchanged.
Contributor
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
FreshData benchmark report —
|
| fixture | n_rows | n_cols | p50 s | p95 s | peak MB | repair % | false-repair % | preserve % | trust | monotonic | export % |
|---|
Authored-code reduction (Metric 6)
kevincostner17
added a commit
that referenced
this pull request
Sep 15, 2026
…ks/; py3.9 aarch64 install note (#427) * docs: regenerate HTML examples for current default output The committed samples predated the report-only duplicate default, the outlier flag default and the not-assessed quality-debt dimensions (#420): rows now read 205 -> 205, revenue shows outliers -> flag (7 flagged), and the quality_debt total is 0.34 with schema_drift/category_churn unknown. Regenerated with scripts/generate_html_examples.py. The action timeline's duration and the quality-debt run_at timestamp are run-dependent and no renderer option omits them, so those lines change on every regeneration. * test: skip benchmark tests when benchmarks/ is absent The sdist ships tests/ but not benchmarks/. conftest.py only filters module-level imports, so tests that read benchmarks/ inside the test body failed or errored from an unpacked sdist. The bench_streaming fixture now uses pytest.importorskip (as the cleanbench suites do), and the entity resolution smoke test and benchmark-stream CLI test skip when the benchmark file is missing. The CLI test also runs from the repo root, since the command resolves benchmarks/ against the working directory. * docs: note Python 3.9 on Linux aarch64 builds the privacy stack from source thinc 8.3.4 and blis 1.2.0 publish no cp39 Linux aarch64 wheel, and spacy 3.8.7 (the py3.9 cap) requires thinc>=8.3.4,<8.4, so no pin avoids the source build. Document it in the install guide and next to the py3.9 caps in pyproject.toml. Dependency pins are unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Two public scores reported results for things they never measured, or misread the data:
fd.evaluate_quality_debtgave a dimension a clean 0.0 when it couldn't be computed, so the gate could saypasswithout having checked it.compute_trust_scoreflagged object columns of lists or dicts as "mixed types", which lowered consistency, while the equivalent nested Arrow column scored 100.Root cause
quality.py: several fallbacks returned(0.0, detail)instead of an unmeasured score:type_instabilitywhen profiling raisedpii_riskwhen the PII scan raised or wasn't availableschema_driftandcategory_churnwhen no baseline was givenEach of these became an assessed item that counted toward the total and was written to the ledger.
enterprise/metrics.py:infer_dtypereturns"mixed"for any object column of containers, even when every value is a list. A nested Arrow list infers as"unknown-array", so the same data scored differently depending on how it was stored.Behaviour change
DebtItem.assessedmechanism from fix: report undetectable duplicates and uniqueness as unknown instead of clean #414. This coverstype_instabilityon profiling failure,pii_riskwhen the scan fails or is unavailable,schema_driftandcategory_churnwithoutbaseline=, and any dimension that wasn't scored. For each such item:scoreandover_thresholdserialise asNone, and the detail says why (exception type only, no message).gate.unassessed, printed as? <dimension>: not assessed — <why>, and shown as n/a in HTML.schema_driftandcategory_churnas not assessed instead of 0.0.baseline=is documented as what enables those dimensions. A run without a baseline also no longer writes a 0.0 that reset drift escalation history.Tests
build_profileanddetect_piipatched to raise, with several exception types includingImportError. The affected dimension is unassessed, and the other eight items match an unpatched run.previous, and the third run escalates tofail.to_dict()items, total, status andsummary()for a fully measured dirty frame with a baseline.list<int64>column. Five genuinely mixed layouts are still flagged.baseline=, so they still assert that duplicates is the only unassessed dimension.Verification
pytest -m "not online and not large": py3.12 / pandas 2.3.3, 6274 passed; py3.9 / pandas 1.5.3, 6245 passed.tests/truthbench: 290 passed.to_dict,summary,to_frame, ledger rows, trust scores) diffed against main on both pandas versions: identical.