Skip to content

fix: report unmeasured quality-debt dimensions as not assessed; list columns aren't mixed types - #420

Merged
kevincostner17 merged 2 commits into
mainfrom
fix/scoring-unassessed-and-list-consistency
Sep 15, 2026
Merged

kevincostner17 merged 2 commits into
mainfrom
fix/scoring-unassessed-and-list-consistency

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

Two public scores reported results for things they never measured, or misread the data:

  • fd.evaluate_quality_debt gave a dimension a clean 0.0 when it couldn't be computed, so the gate could say pass without having checked it.
  • compute_trust_score flagged object columns of lists or dicts as "mixed types", which lowered consistency, while the equivalent nested Arrow column scored 100.

Root cause

  • quality.py: several fallbacks returned (0.0, detail) instead of an unmeasured score:

    • type_instability when profiling raised
    • pii_risk when the PII scan raised or wasn't available
    • schema_drift and category_churn when no baseline was given
    • any dimension missing from the scores

    Each of these became an assessed item that counted toward the total and was written to the ledger.

  • enterprise/metrics.py: infer_dtype returns "mixed" for any object column of containers, even when every value is a list. A nested Arrow list infers as "unknown-array", so the same data scored differently depending on how it was stored.

Behaviour change

  • Unmeasured quality-debt dimensions now use the DebtItem.assessed mechanism from fix: report undetectable duplicates and uniqueness as unknown instead of clean #414. This covers type_instability on profiling failure, pii_risk when the scan fails or is unavailable, schema_drift and category_churn without baseline=, and any dimension that wasn't scored. For each such item:
    • score and over_threshold serialise as None, and the detail says why (exception type only, no message).
    • It's listed in gate.unassessed, printed as ? <dimension>: not assessed — <why>, and shown as n/a in HTML.
    • It counts toward neither the total, the gate status nor the ledger, so escalation keeps comparing against the last run that measured it.
  • Runs with a baseline and a working profile and PII scan are byte-identical to before.
  • Runs without a baseline now report schema_drift and category_churn as not assessed instead of 0.0. baseline= is documented as what enables those dimensions. A run without a baseline also no longer writes a 0.0 that reset drift escalation history.
  • Object columns whose non-null values are all lists, all dicts or all tuples are no longer "mixed types", and score like the nested Arrow equivalent (consistency 66.7 → 100 for the repro frame). Genuinely mixed columns are still flagged: strings with numbers, lists with scalars, lists with dicts or tuples. Plain-frame trust scores are unchanged.

Tests

  • Failures: build_profile and detect_pii patched to raise, with several exception types including ImportError. The affected dimension is unassessed, and the other eight items match an unpatched run.
  • Baseline: without one, drift and churn are unassessed; with one, they're measured. A dimension missing from the scores is unassessed.
  • Ledger: a run without a baseline between two drifting runs records no drift or churn rows and keeps previous, and the third run escalates to fail.
  • Parity: exact to_dict() items, total, status and summary() for a fully measured dirty frame with a baseline.
  • Trust score: uniform lists, empty lists, dicts and tuples are consistent. An object list column matches the Arrow list<int64> column. Five genuinely mixed layouts are still flagged.
  • Existing tests: the fix: report undetectable duplicates and uniqueness as unknown instead of clean #414 tests now pass baseline=, so they still assert that duplicates is the only unassessed dimension.

Verification

  • ruff and mypy clean.
  • pytest -m "not online and not large": py3.12 / pandas 2.3.3, 6274 passed; py3.9 / pandas 1.5.3, 6245 passed.
  • tests/truthbench: 290 passed.
  • Normal-frame outputs (to_dict, summary, to_frame, ledger rows, trust scores) diffed against main on both pandas versions: identical.

evaluate_quality_debt scored a dimension a clean 0.0 when it was never
measured: type_instability when profiling failed, pii_risk when the PII
scan was unavailable or failed, schema_drift and category_churn when no
baseline was passed, and any dimension missing from the scores. Each item
counted as assessed, so the gate could report pass on evidence it never
collected, and the ledger stored a 0.0 that reset escalation history.

These dimensions now use the DebtItem.assessed mechanism from #414: score
and over_threshold serialise as None, the detail says why, the item is
listed in gate.unassessed and shown as "? ... not assessed", and it adds
nothing to the total, the gate status or the ledger. Runs with a baseline
and a working profile and PII scan are byte-identical.
compute_trust_score tagged an object column holding Python lists or dicts
as "mixed types" and lowered consistency (66.7 for one such column out of
three), because infer_dtype returns "mixed" for any object column of
containers. The equivalent nested Arrow list column infers as
"unknown-array" and scored 100.

A column whose non-null values are all lists, all dicts or all tuples is
now uniform and scores like the Arrow equivalent. Columns that mix kinds
(strings with numbers, lists with scalars, lists with dicts or tuples) are
still flagged. The container check only runs when infer_dtype already
reports mixed, so scores for plain frames are unchanged.
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 27606f3b-658e-46ac-9912-00a138105c15


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit 757a24e into main Sep 15, 2026
22 checks passed
kevincostner17 added a commit that referenced this pull request Sep 15, 2026
…ks/; py3.9 aarch64 install note (#427)

* docs: regenerate HTML examples for current default output

The committed samples predated the report-only duplicate default, the
outlier flag default and the not-assessed quality-debt dimensions (#420):
rows now read 205 -> 205, revenue shows outliers -> flag (7 flagged), and
the quality_debt total is 0.34 with schema_drift/category_churn unknown.

Regenerated with scripts/generate_html_examples.py. The action timeline's
duration and the quality-debt run_at timestamp are run-dependent and no
renderer option omits them, so those lines change on every regeneration.

* test: skip benchmark tests when benchmarks/ is absent

The sdist ships tests/ but not benchmarks/. conftest.py only filters
module-level imports, so tests that read benchmarks/ inside the test body
failed or errored from an unpacked sdist. The bench_streaming fixture now
uses pytest.importorskip (as the cleanbench suites do), and the entity
resolution smoke test and benchmark-stream CLI test skip when the
benchmark file is missing. The CLI test also runs from the repo root, since
the command resolves benchmarks/ against the working directory.

* docs: note Python 3.9 on Linux aarch64 builds the privacy stack from source

thinc 8.3.4 and blis 1.2.0 publish no cp39 Linux aarch64 wheel, and spacy
3.8.7 (the py3.9 cap) requires thinc>=8.3.4,<8.4, so no pin avoids the
source build. Document it in the install guide and next to the py3.9 caps
in pyproject.toml. Dependency pins are unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant