fix(enterprise): handle nested Arrow columns in trust scoring - #413
Merged
Merged
Conversation
compute_trust_score, and therefore clean_enterprise, crashed with ArrowNotImplementedError on a frame with a nested Arrow column (list, large_list, struct or map). _is_constant calls Series.nunique and the uniqueness dimension calls DataFrame.duplicated; for nested Arrow dtypes pyarrow has no unique/dictionary_encode kernel and raises ArrowNotImplementedError, a NotImplementedError subclass. Both guards only caught TypeError, which is what object columns holding lists raise. Catch NotImplementedError alongside TypeError so nested Arrow columns follow the existing unhashable-cell path: the column is not reported as constant and uniqueness falls back to 100. Scores for frames without nested columns are unchanged.
Contributor
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
FreshData benchmark report —
|
| fixture | n_rows | n_cols | p50 s | p95 s | peak MB | repair % | false-repair % | preserve % | trust | monotonic | export % |
|---|
Authored-code reduction (Metric 6)
kevincostner17
added a commit
that referenced
this pull request
Sep 15, 2026
… of clean (#414) * fix(quality): report undetectable duplicates as unknown instead of clean _score_debt counted remaining duplicates with cleaned.duplicated() and, when that raised (TypeError for object columns of lists or dicts, and since #412 NotImplementedError for nested Arrow list/struct/map columns), set dup_remaining = 0. "Cannot check" was therefore scored as "no duplicates": the dimension read 0.0 with "0 duplicate row(s) detected", counted as clean evidence toward the gate, and was written to the ledger as a real 0.0, so a later measured run could be compared against a score that was never measured. DebtItem gains assessed (default True). When detection is impossible and nothing was removed, the duplicates dimension is unassessed: it is never over threshold or worsening, adds nothing to total_score, never drives the gate status, is skipped when writing the ledger (escalation compares against the last run that measured it), serialises with score and over_threshold as None, and is listed in QualityDebtGate.unassessed and in a "? duplicates: not assessed" summary line. The detail names the columns that block detection. If rows were already removed through a hashable duplicate_subset, that count stays a known lower bound. Frames where duplicates can be detected take the same code path and produce identical scores, dicts, frames and summaries. * fix(enterprise): report undetectable uniqueness as unknown in trust scores compute_trust_score computed uniqueness from frame.duplicated() and, when that raised (TypeError for object columns of lists or dicts, NotImplementedError for nested Arrow list/struct/map columns, caught since #413), set uniqueness = 100.0. "Cannot check" was scored as "no duplicates": a perfect dimension blended into overall at its full weight, lifting the trust score of frames whose duplicates were never checked. It also disagreed with the quality-debt duplicates dimension, which now reports the same situation as not assessed. When duplicate rows cannot be checked, uniqueness is now NaN (TrustScore.uniqueness_assessed is False) and overall blends completeness, validity and consistency with their weights renormalised to sum to 1 (equal weights if all of the weight was on uniqueness). to_dict() gives uniqueness None, str() and both Markdown tables show n/a, and each column that blocks the check carries an "unhashable values: duplicate rows not checked" issue. Frames where duplicates can be detected keep the original four-way expression and produce identical scores, dicts and renderings. * fix(render): show unknown trust uniqueness as n/a in the report summary TrustScore.uniqueness is NaN when duplicate rows could not be checked. The report summary formatted it with :.0f and printed 'uniqueness nan'; it now prints 'uniqueness n/a', matching TrustScore's own text and Markdown output. * fix(quality): compare debt against each dimension's last recorded score Unassessed dimensions are no longer written to the ledger, but _previous_scores read only the latest run's items. A run after an unassessed one therefore lost its previous score and could not be marked worsening. Read each dimension's score from the last run that recorded it; when every run records every dimension the result is unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
compute_trust_score, and thereforeclean_enterprise, crashed on any frame with a nested Arrow column (list, large_list, struct or map), raisingpyarrow.lib.ArrowNotImplementedError: Function 'unique' has no kernel matching input types (list<item: string>). Trust scoring now treats these columns the same way it already treats object columns holding lists: the value can't be computed, so it doesn't crash. The column isn't reported as constant, and uniqueness falls back to 100. Scores for frames without nested columns are unchanged.This follows the merged #412, which fixed the same kind of crash in
profile.py,steps/duplicates.py,engine/context.pyandquality.py. With both changes,clean_enterpriseworks end to end on nested Arrow columns.Root cause
Two calls in
enterprise/metrics.pycan't handle nested Arrow data:_is_constantcallsSeries.nunique.DataFrame.duplicated.pyarrow has no
uniqueordictionary_encodekernel for nested types, so both calls raiseArrowNotImplementedError, a subclass ofNotImplementedError. The guards around them only caughtTypeError, which is what object columns holding lists raise. Both guards now catchNotImplementedErrortoo.Tests
New
tests/test_enterprise_metrics_nested.py:clean_enterpriseend to end for all four nested kindsNested-Arrow tests skip when pyarrow or
pd.ArrowDtypeis unavailable.Verification
metrics.py, the nested-Arrow tests fail and the parity and object-list tests pass.ruff check .: clean.mypy src/freshdata: clean.pytest -m "not online and not large":