Skip to content

fix(profile): handle nested Arrow and list-valued columns in duplicate detection - #412

Merged
kevincostner17 merged 1 commit into
mainfrom
fix/profile-arrow-nested-columns
Sep 15, 2026
Merged

kevincostner17 merged 1 commit into
mainfrom
fix/profile-arrow-nested-columns

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

fd.profile and fd.clean crashed with pyarrow.lib.ArrowNotImplementedError on any frame with a nested Arrow column (pd.ArrowDtype list, large_list, struct or map). Nested Arrow columns now get the same handling as object columns holding Python lists or dicts:

  • the profile reports duplicate_rows=None and unique=None
  • the clean step skips duplicate detection with a report note instead of crashing

On pandas 1.5 this also fixes a second crash. Removing duplicates with a duplicate_subset that excludes the nested column failed with ValueError for duplicate_keep = first, last or drop.

Root cause

DataFrame.duplicated, Series.nunique and Series.unique raise TypeError for object columns of lists or dicts. The existing guards catch that as "unhashable, detection impossible". Nested Arrow columns go through ArrowExtensionArray, whose dictionary_encode and unique kernels don't support these types. They raise ArrowNotImplementedError, a NotImplementedError subclass, so the error escaped every except TypeError guard.

Separately, _filter_rows rebuilt each column through to_numpy(). On pandas 1.5 that raises for Arrow lists of different lengths, and returns a 2-D array for equal-length lists and maps.

The fix:

  • The guards in profile.py, steps/duplicates.py, engine/context.py and quality.py now also catch NotImplementedError.
  • _filter_rows takes such columns positionally from their extension array.
  • Columns that already converted to 1-D numpy arrays keep the existing path, so output for frames without nested columns is unchanged.

Tests

New tests/test_nested_values.py:

  • The original list<string> repro, through fd.profile, str() and to_dict().
  • list, equal-length list, large_list, struct and map Arrow columns through fd.profile and fd.clean, with and without drop_duplicates.
  • Rows that match or differ only in the nested column: duplicate_rows is None, every row is kept, and the skip is reported.
  • A duplicate_subset that excludes the nested column, for every duplicate_keep mode: the exact rows are kept, the removed count is correct, and the nested dtype is preserved.
  • Object columns holding lists or dicts.
  • Nested Arrow profiles match object-list profiles.
  • Parity: plain frames and Arrow scalar frames keep their exact duplicate and unique counts.

Arrow tests use pytest.importorskip("pyarrow") and skip when pd.ArrowDtype is unavailable. 47 of the new tests fail on the previous source.

Verification

  • ruff check .: passed
  • mypy src/freshdata: no issues
  • pytest -m "not online and not large" on Python 3.12 / pandas 2.3.3 / pyarrow 25: 5947 passed, 14 skipped
  • Same on Python 3.9 / pandas 1.5.3 / pyarrow 21: 5918 passed, 18 skipped

…e detection

fd.profile and fd.clean crashed with pyarrow.lib.ArrowNotImplementedError
on frames holding a nested Arrow column (pd.ArrowDtype list, large_list,
struct or map).

Root cause: DataFrame.duplicated factorizes every column. For object
columns holding Python lists or dicts that raises TypeError, which the
duplicate guards already treat as "cells are unhashable, detection is
impossible": the profile reports duplicate_rows=None and the clean step
skips with a report note. Nested Arrow columns factorize through
ArrowExtensionArray, whose dictionary_encode/unique kernels do not exist
for those types, so they raise ArrowNotImplementedError instead. That is
a NotImplementedError subclass, not a TypeError, so it escaped every
guard. Series.nunique/unique in the per-column profile hit the same error.

The guards in profile.py (duplicate scan, _safe_nunique, _sample_values),
steps/duplicates.py, engine/context.py and quality.py now also catch
NotImplementedError, so nested Arrow columns get exactly the treatment
object list columns already had. Frames without such columns take the
same code path as before.

A second crash surfaced on pandas 1.5 once detection worked: with a
duplicate_subset that excludes the nested column, removing rows went
through _filter_rows, which rebuilds each column via to_numpy(). pandas
1.5 cannot turn ragged Arrow lists into a 1-D numpy array (ValueError)
and silently returns a 2-D array for equal-length lists and maps. Such
columns are now taken positionally from their extension array; every
other column keeps the numpy path.
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 3a916899-9bf9-4895-8468-b247c2e567ad


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit b492a44 into main Sep 15, 2026
22 checks passed
kevincostner17 added a commit that referenced this pull request Sep 15, 2026
… of clean (#414)

* fix(quality): report undetectable duplicates as unknown instead of clean

_score_debt counted remaining duplicates with cleaned.duplicated() and,
when that raised (TypeError for object columns of lists or dicts, and
since #412 NotImplementedError for nested Arrow list/struct/map
columns), set dup_remaining = 0. "Cannot check" was therefore scored as
"no duplicates": the dimension read 0.0 with "0 duplicate row(s)
detected", counted as clean evidence toward the gate, and was written
to the ledger as a real 0.0, so a later measured run could be compared
against a score that was never measured.

DebtItem gains assessed (default True). When detection is impossible
and nothing was removed, the duplicates dimension is unassessed: it is
never over threshold or worsening, adds nothing to total_score, never
drives the gate status, is skipped when writing the ledger (escalation
compares against the last run that measured it), serialises with score
and over_threshold as None, and is listed in QualityDebtGate.unassessed
and in a "? duplicates: not assessed" summary line. The detail names the
columns that block detection. If rows were already removed through a
hashable duplicate_subset, that count stays a known lower bound.

Frames where duplicates can be detected take the same code path and
produce identical scores, dicts, frames and summaries.

* fix(enterprise): report undetectable uniqueness as unknown in trust scores

compute_trust_score computed uniqueness from frame.duplicated() and,
when that raised (TypeError for object columns of lists or dicts,
NotImplementedError for nested Arrow list/struct/map columns, caught
since #413), set uniqueness = 100.0. "Cannot check" was scored as "no
duplicates": a perfect dimension blended into overall at its full
weight, lifting the trust score of frames whose duplicates were never
checked. It also disagreed with the quality-debt duplicates dimension,
which now reports the same situation as not assessed.

When duplicate rows cannot be checked, uniqueness is now NaN
(TrustScore.uniqueness_assessed is False) and overall blends
completeness, validity and consistency with their weights renormalised
to sum to 1 (equal weights if all of the weight was on uniqueness).
to_dict() gives uniqueness None, str() and both Markdown tables show
n/a, and each column that blocks the check carries an "unhashable
values: duplicate rows not checked" issue.

Frames where duplicates can be detected keep the original four-way
expression and produce identical scores, dicts and renderings.

* fix(render): show unknown trust uniqueness as n/a in the report summary

TrustScore.uniqueness is NaN when duplicate rows could not be checked. The report summary formatted it with :.0f and printed 'uniqueness nan'; it now prints 'uniqueness n/a', matching TrustScore's own text and Markdown output.

* fix(quality): compare debt against each dimension's last recorded score

Unassessed dimensions are no longer written to the ledger, but _previous_scores read only the latest run's items. A run after an unassessed one therefore lost its previous score and could not be marked worsening. Read each dimension's score from the last run that recorded it; when every run records every dimension the result is unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant