fix: report undetectable duplicates and uniqueness as unknown instead of clean - #414
Merged
Merged
Conversation
_score_debt counted remaining duplicates with cleaned.duplicated() and, when that raised (TypeError for object columns of lists or dicts, and since #412 NotImplementedError for nested Arrow list/struct/map columns), set dup_remaining = 0. "Cannot check" was therefore scored as "no duplicates": the dimension read 0.0 with "0 duplicate row(s) detected", counted as clean evidence toward the gate, and was written to the ledger as a real 0.0, so a later measured run could be compared against a score that was never measured. DebtItem gains assessed (default True). When detection is impossible and nothing was removed, the duplicates dimension is unassessed: it is never over threshold or worsening, adds nothing to total_score, never drives the gate status, is skipped when writing the ledger (escalation compares against the last run that measured it), serialises with score and over_threshold as None, and is listed in QualityDebtGate.unassessed and in a "? duplicates: not assessed" summary line. The detail names the columns that block detection. If rows were already removed through a hashable duplicate_subset, that count stays a known lower bound. Frames where duplicates can be detected take the same code path and produce identical scores, dicts, frames and summaries.
…cores compute_trust_score computed uniqueness from frame.duplicated() and, when that raised (TypeError for object columns of lists or dicts, NotImplementedError for nested Arrow list/struct/map columns, caught since #413), set uniqueness = 100.0. "Cannot check" was scored as "no duplicates": a perfect dimension blended into overall at its full weight, lifting the trust score of frames whose duplicates were never checked. It also disagreed with the quality-debt duplicates dimension, which now reports the same situation as not assessed. When duplicate rows cannot be checked, uniqueness is now NaN (TrustScore.uniqueness_assessed is False) and overall blends completeness, validity and consistency with their weights renormalised to sum to 1 (equal weights if all of the weight was on uniqueness). to_dict() gives uniqueness None, str() and both Markdown tables show n/a, and each column that blocks the check carries an "unhashable values: duplicate rows not checked" issue. Frames where duplicates can be detected keep the original four-way expression and produce identical scores, dicts and renderings.
TrustScore.uniqueness is NaN when duplicate rows could not be checked. The report summary formatted it with :.0f and printed 'uniqueness nan'; it now prints 'uniqueness n/a', matching TrustScore's own text and Markdown output.
Contributor
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
FreshData benchmark report —
|
| fixture | n_rows | n_cols | p50 s | p95 s | peak MB | repair % | false-repair % | preserve % | trust | monotonic | export % |
|---|
Authored-code reduction (Metric 6)
Unassessed dimensions are no longer written to the ledger, but _previous_scores read only the latest run's items. A run after an unassessed one therefore lost its previous score and could not be marked worsening. Read each dimension's score from the last run that recorded it; when every run records every dimension the result is unchanged.
kevincostner17
added a commit
that referenced
this pull request
Sep 15, 2026
…columns aren't mixed types (#420) * fix(quality): report unmeasured quality-debt dimensions as not assessed evaluate_quality_debt scored a dimension a clean 0.0 when it was never measured: type_instability when profiling failed, pii_risk when the PII scan was unavailable or failed, schema_drift and category_churn when no baseline was passed, and any dimension missing from the scores. Each item counted as assessed, so the gate could report pass on evidence it never collected, and the ledger stored a 0.0 that reset escalation history. These dimensions now use the DebtItem.assessed mechanism from #414: score and over_threshold serialise as None, the detail says why, the item is listed in gate.unassessed and shown as "? ... not assessed", and it adds nothing to the total, the gate status or the ledger. Runs with a baseline and a working profile and PII scan are byte-identical. * fix(enterprise): do not flag uniform list or dict columns as mixed types compute_trust_score tagged an object column holding Python lists or dicts as "mixed types" and lowered consistency (66.7 for one such column out of three), because infer_dtype returns "mixed" for any object column of containers. The equivalent nested Arrow list column infers as "unknown-array" and scored 100. A column whose non-null values are all lists, all dicts or all tuples is now uniform and scores like the Arrow equivalent. Columns that mix kinds (strings with numbers, lists with scalars, lists with dicts or tuples) are still flagged. The container check only runs when infer_dtype already reports mixed, so scores for plain frames are unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Duplicate rows can't be checked when a column holds unhashable values: object columns of lists or dicts, or nested Arrow list/struct/map columns. In that case, both the quality-debt
duplicatesdimension and the trust-scoreuniquenessdimension used to report a clean result. They are now reported as unknown: excluded from totals, weights and the gate, and labelled as unknown wherever they are shown.Related: #412 and #413 (catch the
NotImplementedErrorfrom nested Arrow columns), and #264 / #286 (introduced remaining-duplicate scoring for the debt dimension).Root cause
_score_debtcounts remaining duplicates withcleaned.duplicated(). When that raised, it setdup_remaining = 0. The error isTypeErrorfor lists and dicts, and since fix(profile): handle nested Arrow and list-valued columns in duplicate detection #412 alsoNotImplementedErrorfor nested Arrow columns.compute_trust_scoredid the same withuniqueness = 100.0, which raisedoverallby the dimension's full weight.Behaviour change
Quality-debt report (
DebtItem.assessed, defaultTrue). When duplicates can't be checked and no rows were removed, the item:total_score, and never changes the gate statusscoreandover_thresholdasNoneQualityDebtGate.unassessed, printed as? duplicates: not assessed — …insummary(), and shown asn/ain HTMLIf rows were already removed through a hashable
duplicate_subset, that count is kept as a known lower bound.Trust score.
TrustScore.uniquenessisNaNanduniqueness_assessedisFalse. It'sNoneinto_dict(), andn/ainstr(), both Markdown tables and the report summary.overallblends completeness, validity and consistency with their weights renormalised, or equal weights if all the weight was on uniqueness.unhashable values: duplicate rows not checkedissue.Frames where duplicates can be checked produce identical scores, dicts, frames, summaries and renderings.
Example, a frame with a list column and a real duplicate row:
score 0.0, "0 duplicate row(s) detected"score null, "duplicate rows could not be checked: column(s) 'tags' hold unhashable values"Tests
pd.ArrowDtype). The duplicates dimension is:warnedand from the totalto_dict→ JSONn/ain HTML andto_framewarn,failandwarn_then_failduplicate_subsetremoved rows, the lower bound is kept.to_dictoutput for plain frames with and without duplicates.Nonein JSON andn/ain text, Markdown (includingQualityReport) and the report summary, and the column issue is recorded.clean_enterpriseJSON.Verification
ruff check .: all checks passedmypy src/freshdata: no issues in 204 source filespytest -m "not online and not large"on Python 3.12: 5976 passed, 14 skippedpytest tests/truthbench: 265 passedmain.