fix(profile): handle nested Arrow and list-valued columns in duplicate detection - #412
Merged
Merged
Conversation
…e detection fd.profile and fd.clean crashed with pyarrow.lib.ArrowNotImplementedError on frames holding a nested Arrow column (pd.ArrowDtype list, large_list, struct or map). Root cause: DataFrame.duplicated factorizes every column. For object columns holding Python lists or dicts that raises TypeError, which the duplicate guards already treat as "cells are unhashable, detection is impossible": the profile reports duplicate_rows=None and the clean step skips with a report note. Nested Arrow columns factorize through ArrowExtensionArray, whose dictionary_encode/unique kernels do not exist for those types, so they raise ArrowNotImplementedError instead. That is a NotImplementedError subclass, not a TypeError, so it escaped every guard. Series.nunique/unique in the per-column profile hit the same error. The guards in profile.py (duplicate scan, _safe_nunique, _sample_values), steps/duplicates.py, engine/context.py and quality.py now also catch NotImplementedError, so nested Arrow columns get exactly the treatment object list columns already had. Frames without such columns take the same code path as before. A second crash surfaced on pandas 1.5 once detection worked: with a duplicate_subset that excludes the nested column, removing rows went through _filter_rows, which rebuilds each column via to_numpy(). pandas 1.5 cannot turn ragged Arrow lists into a 1-D numpy array (ValueError) and silently returns a 2-D array for equal-length lists and maps. Such columns are now taken positionally from their extension array; every other column keeps the numpy path.
Contributor
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
FreshData benchmark report —
|
| fixture | n_rows | n_cols | p50 s | p95 s | peak MB | repair % | false-repair % | preserve % | trust | monotonic | export % |
|---|
Authored-code reduction (Metric 6)
This was referenced Sep 15, 2026
kevincostner17
added a commit
that referenced
this pull request
Sep 15, 2026
… of clean (#414) * fix(quality): report undetectable duplicates as unknown instead of clean _score_debt counted remaining duplicates with cleaned.duplicated() and, when that raised (TypeError for object columns of lists or dicts, and since #412 NotImplementedError for nested Arrow list/struct/map columns), set dup_remaining = 0. "Cannot check" was therefore scored as "no duplicates": the dimension read 0.0 with "0 duplicate row(s) detected", counted as clean evidence toward the gate, and was written to the ledger as a real 0.0, so a later measured run could be compared against a score that was never measured. DebtItem gains assessed (default True). When detection is impossible and nothing was removed, the duplicates dimension is unassessed: it is never over threshold or worsening, adds nothing to total_score, never drives the gate status, is skipped when writing the ledger (escalation compares against the last run that measured it), serialises with score and over_threshold as None, and is listed in QualityDebtGate.unassessed and in a "? duplicates: not assessed" summary line. The detail names the columns that block detection. If rows were already removed through a hashable duplicate_subset, that count stays a known lower bound. Frames where duplicates can be detected take the same code path and produce identical scores, dicts, frames and summaries. * fix(enterprise): report undetectable uniqueness as unknown in trust scores compute_trust_score computed uniqueness from frame.duplicated() and, when that raised (TypeError for object columns of lists or dicts, NotImplementedError for nested Arrow list/struct/map columns, caught since #413), set uniqueness = 100.0. "Cannot check" was scored as "no duplicates": a perfect dimension blended into overall at its full weight, lifting the trust score of frames whose duplicates were never checked. It also disagreed with the quality-debt duplicates dimension, which now reports the same situation as not assessed. When duplicate rows cannot be checked, uniqueness is now NaN (TrustScore.uniqueness_assessed is False) and overall blends completeness, validity and consistency with their weights renormalised to sum to 1 (equal weights if all of the weight was on uniqueness). to_dict() gives uniqueness None, str() and both Markdown tables show n/a, and each column that blocks the check carries an "unhashable values: duplicate rows not checked" issue. Frames where duplicates can be detected keep the original four-way expression and produce identical scores, dicts and renderings. * fix(render): show unknown trust uniqueness as n/a in the report summary TrustScore.uniqueness is NaN when duplicate rows could not be checked. The report summary formatted it with :.0f and printed 'uniqueness nan'; it now prints 'uniqueness n/a', matching TrustScore's own text and Markdown output. * fix(quality): compare debt against each dimension's last recorded score Unassessed dimensions are no longer written to the ledger, but _previous_scores read only the latest run's items. A run after an unassessed one therefore lost its previous score and could not be marked worsening. Read each dimension's score from the last run that recorded it; when every run records every dimension the result is unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
fd.profileandfd.cleancrashed withpyarrow.lib.ArrowNotImplementedErroron any frame with a nested Arrow column (pd.ArrowDtypelist, large_list, struct or map). Nested Arrow columns now get the same handling as object columns holding Python lists or dicts:duplicate_rows=Noneandunique=NoneOn pandas 1.5 this also fixes a second crash. Removing duplicates with a
duplicate_subsetthat excludes the nested column failed withValueErrorforduplicate_keep=first,lastordrop.Root cause
DataFrame.duplicated,Series.nuniqueandSeries.uniqueraiseTypeErrorfor object columns of lists or dicts. The existing guards catch that as "unhashable, detection impossible". Nested Arrow columns go throughArrowExtensionArray, whosedictionary_encodeanduniquekernels don't support these types. They raiseArrowNotImplementedError, aNotImplementedErrorsubclass, so the error escaped everyexcept TypeErrorguard.Separately,
_filter_rowsrebuilt each column throughto_numpy(). On pandas 1.5 that raises for Arrow lists of different lengths, and returns a 2-D array for equal-length lists and maps.The fix:
profile.py,steps/duplicates.py,engine/context.pyandquality.pynow also catchNotImplementedError._filter_rowstakes such columns positionally from their extension array.Tests
New
tests/test_nested_values.py:list<string>repro, throughfd.profile,str()andto_dict().fd.profileandfd.clean, with and withoutdrop_duplicates.duplicate_rowsisNone, every row is kept, and the skip is reported.duplicate_subsetthat excludes the nested column, for everyduplicate_keepmode: the exact rows are kept, the removed count is correct, and the nested dtype is preserved.Arrow tests use
pytest.importorskip("pyarrow")and skip whenpd.ArrowDtypeis unavailable. 47 of the new tests fail on the previous source.Verification
ruff check .: passedmypy src/freshdata: no issuespytest -m "not online and not large"on Python 3.12 / pandas 2.3.3 / pyarrow 25: 5947 passed, 14 skipped