Skip to content

fix(enterprise): handle nested Arrow columns in trust scoring - #413

Merged
kevincostner17 merged 1 commit into
mainfrom
fix/enterprise-metrics-nested-columns
Sep 15, 2026
Merged

kevincostner17 merged 1 commit into
mainfrom
fix/enterprise-metrics-nested-columns

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

compute_trust_score, and therefore clean_enterprise, crashed on any frame with a nested Arrow column (list, large_list, struct or map), raising pyarrow.lib.ArrowNotImplementedError: Function 'unique' has no kernel matching input types (list<item: string>). Trust scoring now treats these columns the same way it already treats object columns holding lists: the value can't be computed, so it doesn't crash. The column isn't reported as constant, and uniqueness falls back to 100. Scores for frames without nested columns are unchanged.

This follows the merged #412, which fixed the same kind of crash in profile.py, steps/duplicates.py, engine/context.py and quality.py. With both changes, clean_enterprise works end to end on nested Arrow columns.

Root cause

Two calls in enterprise/metrics.py can't handle nested Arrow data:

  • _is_constant calls Series.nunique.
  • The uniqueness dimension calls DataFrame.duplicated.

pyarrow has no unique or dictionary_encode kernel for nested types, so both calls raise ArrowNotImplementedError, a subclass of NotImplementedError. The guards around them only caught TypeError, which is what object columns holding lists raise. Both guards now catch NotImplementedError too.

Tests

New tests/test_enterprise_metrics_nested.py:

  • the list repro, plus trust scoring for list, large_list, struct and map columns: no crash, uniqueness 100, not flagged constant, input frame not modified
  • identical nested values aren't reported as constant
  • an object column holding lists gets the same fallback
  • parity: exact dimension values and per-column issues for a plain frame
  • clean_enterprise end to end for all four nested kinds

Nested-Arrow tests skip when pyarrow or pd.ArrowDtype is unavailable.

Verification

  • Against the previous metrics.py, the nested-Arrow tests fail and the parity and object-list tests pass.
  • ruff check .: clean. mypy src/freshdata: clean.
  • pytest -m "not online and not large":
    • Python 3.12 (pandas 2.3.3, pyarrow 25.0.1): 5962 passed, 14 skipped
    • Python 3.9 (pandas 1.5.3, pyarrow 21.0.0): 5933 passed, 18 skipped

compute_trust_score, and therefore clean_enterprise, crashed with
ArrowNotImplementedError on a frame with a nested Arrow column (list,
large_list, struct or map). _is_constant calls Series.nunique and the
uniqueness dimension calls DataFrame.duplicated; for nested Arrow dtypes
pyarrow has no unique/dictionary_encode kernel and raises
ArrowNotImplementedError, a NotImplementedError subclass. Both guards
only caught TypeError, which is what object columns holding lists raise.

Catch NotImplementedError alongside TypeError so nested Arrow columns
follow the existing unhashable-cell path: the column is not reported as
constant and uniqueness falls back to 100. Scores for frames without
nested columns are unchanged.
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: dabfcbe3-525c-4132-b71c-8ea8728c86a0


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit 4af1cb8 into main Sep 15, 2026
22 checks passed
kevincostner17 added a commit that referenced this pull request Sep 15, 2026
… of clean (#414)

* fix(quality): report undetectable duplicates as unknown instead of clean

_score_debt counted remaining duplicates with cleaned.duplicated() and,
when that raised (TypeError for object columns of lists or dicts, and
since #412 NotImplementedError for nested Arrow list/struct/map
columns), set dup_remaining = 0. "Cannot check" was therefore scored as
"no duplicates": the dimension read 0.0 with "0 duplicate row(s)
detected", counted as clean evidence toward the gate, and was written
to the ledger as a real 0.0, so a later measured run could be compared
against a score that was never measured.

DebtItem gains assessed (default True). When detection is impossible
and nothing was removed, the duplicates dimension is unassessed: it is
never over threshold or worsening, adds nothing to total_score, never
drives the gate status, is skipped when writing the ledger (escalation
compares against the last run that measured it), serialises with score
and over_threshold as None, and is listed in QualityDebtGate.unassessed
and in a "? duplicates: not assessed" summary line. The detail names the
columns that block detection. If rows were already removed through a
hashable duplicate_subset, that count stays a known lower bound.

Frames where duplicates can be detected take the same code path and
produce identical scores, dicts, frames and summaries.

* fix(enterprise): report undetectable uniqueness as unknown in trust scores

compute_trust_score computed uniqueness from frame.duplicated() and,
when that raised (TypeError for object columns of lists or dicts,
NotImplementedError for nested Arrow list/struct/map columns, caught
since #413), set uniqueness = 100.0. "Cannot check" was scored as "no
duplicates": a perfect dimension blended into overall at its full
weight, lifting the trust score of frames whose duplicates were never
checked. It also disagreed with the quality-debt duplicates dimension,
which now reports the same situation as not assessed.

When duplicate rows cannot be checked, uniqueness is now NaN
(TrustScore.uniqueness_assessed is False) and overall blends
completeness, validity and consistency with their weights renormalised
to sum to 1 (equal weights if all of the weight was on uniqueness).
to_dict() gives uniqueness None, str() and both Markdown tables show
n/a, and each column that blocks the check carries an "unhashable
values: duplicate rows not checked" issue.

Frames where duplicates can be detected keep the original four-way
expression and produce identical scores, dicts and renderings.

* fix(render): show unknown trust uniqueness as n/a in the report summary

TrustScore.uniqueness is NaN when duplicate rows could not be checked. The report summary formatted it with :.0f and printed 'uniqueness nan'; it now prints 'uniqueness n/a', matching TrustScore's own text and Markdown output.

* fix(quality): compare debt against each dimension's last recorded score

Unassessed dimensions are no longer written to the ledger, but _previous_scores read only the latest run's items. A run after an unassessed one therefore lost its previous score and could not be marked worsening. Read each dimension's score from the last run that recorded it; when every run records every dimension the result is unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant