Skip to content

fix(learning): learn from categorical columns whose categories differ - #410

Merged
kevincostner17 merged 1 commit into
mainfrom
fix/learn-categorical-categories
Sep 15, 2026
Merged

kevincostner17 merged 1 commit into
mainfrom
fix/learn-categorical-categories

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

fd.learn crashed when a column was categorical in both the messy and clean frames but with different categories. This included the same categories in a different order, and ordered vs unordered categoricals. The error was TypeError: Categoricals can only be compared if 'categories' are the same. Every fd.learn path hit it: positional and keyed alignment, both privacy modes, and context=. Categorical columns are now compared and learned by their values, and the learned profile is the same as one learned from the object-dtype equivalent.

Root cause

The diff stage (learning/diff.py::_column_diffs) finds changed cells with an elementwise messy.eq(clean). pandas only compares two Categoricals whose categories are identical. A messy/clean pair rarely has identical categories, because fixing the values is exactly what changes them. The later stages (classify, extract, holdout evaluation) already handled categorical values. Pairs with matching categories, or with a categorical on only one side, did not crash and already matched the object-dtype profile.

Fix

_column_diffs now converts a categorical side to its object values before comparing, which is what the object-dtype column would hold. dtype_changes still reports the original dtypes, and non-categorical columns are untouched.

Tests

New tests/learning/test_categorical.py (14 of its tests fail without the fix):

  • the original repro
  • the diff stage compares categoricals by value and reports no dtype change
  • learning gives the same profile as an object-dtype run, with and without key=, for:
    • both categorical, with different or the same categories
    • categorical messy with object clean, and the reverse
    • both ordered
    • ordered with a different category order
    • ordered vs unordered
  • a categorical key column whose categories are in a different order
  • missing cells: a cell missing on both sides is not a diff, and a placeholder that becomes missing is learned
  • filled-in categorical cells never become a value map
  • a sensitive categorical column learns like object under privacy="none", and its masked flags match under privacy="mask"
  • save/load round trip (via tmp_path) keeps profile_id, and replaying with fd.clean(profile=...) on a categorical batch gives the same values as the object-dtype profile on an object batch

Verification

  • ruff check .: clean
  • mypy src/freshdata: no issues
  • pytest -m "not online and not large", Python 3.12 / pandas 2.3.3: 5660 passed, 14 skipped
  • the same on Python 3.9 / pandas 1.5.3: 5656 passed, 18 skipped

fd.learn raised "TypeError: Categoricals can only be compared if
'categories' are the same" whenever a shared column was categorical on
both sides with different categories (including the same categories in a
different order, or differing orderedness).

Root cause: the diff stage (learning/diff.py::_column_diffs) finds changed
cells with an elementwise `messy.eq(clean)`. pandas only compares two
Categoricals when their categories are identical, and a messy/clean pair
almost never shares categories, because repairing the values is exactly
what changes them. Every fd.learn path runs this comparison (positional
and keyed alignment, privacy="mask"/"none", context=), so they all
crashed before classification. Pairs that happened to share categories,
or mixed categorical with object, compared fine and already learned the
same profile as the object-dtype run.

Fix: cells are compared by value, so the diff stage decodes a categorical
side to its object values before comparing. That is exactly the
object-dtype equivalent. Every later stage already handled categorical
values, so after the diff the learned rules, value maps, examples,
embedded memory and audit match the object-dtype run. The recorded
dtype_changes still use the original dtypes. Non-categorical columns do
not go through the new branch, so their output is unchanged.
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: fe9d32c2-905b-464e-a9cc-a381384d4551


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit dc8ac5e into main Sep 15, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant