Skip to content

fix(learning): text dtypes compatible in replay; merge list memory patterns - #419

Merged
kevincostner17 merged 3 commits into
mainfrom
fix/learning-replay-dtype-and-merge
Sep 15, 2026
Merged

kevincostner17 merged 3 commits into
mainfrom
fix/learning-replay-dtype-and-merge

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

This fixes two bugs in learned profiles and documents one existing behaviour.

  • Replay across text dtypes: replaying a profile on the same values in a different text dtype (object, string, Arrow string or categorical) no longer reports drift and drops that column's value maps.
  • Merge crash: merge with union_min_precision or error_on_conflict no longer crashes when the embedded memory holds a semantic_repairs list.
  • Docs: the privacy module docstring now explains that masked tokens are per profile. The salt comes from the frame signature, so merged profiles can't match each other's masked entries.

Root cause

Replay: check_profile_drift treated only object as text. When a profile learned on a categorical or string column was replayed on object data (or the reverse), it reported mild drift and didn't replay that column's value maps and hints. The same check also missed drift when a text categorical became numeric.

Merge: the union path assumed every value_patterns entry was a {raw: clean} mapping. It called dict() on the semantic_repairs list, which raised ValueError: dictionary update sequence element #0 has length 4. error_on_conflict goes through the same path, and memory was only merged after the conflict check.

Changes:

  • The replay drift check now asks whether the frame's dtype and the stored learned dtype are text, using the existing is_text_dtype helper. Text dtypes are interchangeable, and text vs non-text still counts as drift. A stored category carries no category info, so it is treated as text, as memory signatures already do.
  • The memory merge now unions list patterns and removes duplicates. When both sides propose different values for the same repair:
    • under union_min_precision, the repair is dropped and recorded;
    • under error_on_conflict, it raises ProfileMergeError.
      Scalar patterns that differ are handled the same way. Conflict messages name only the column and issue type, never raw values. prefer_self and prefer_other are unchanged.

Behaviour change: a text categorical replayed as numeric is now reported as drift. So is a numeric categorical learned and then replayed as int64.

Tests

  • Replay:
    • learned on object and replayed on category, string, Arrow string and Arrow large_string, and each reverse direction: severity none, every column compatible, value maps replayed;
    • text to numeric still drifts;
    • a numeric categorical is not treated as text.
  • Merge, under all four strategies:
    • semantic_repairs union with duplicates removed;
    • conflicting repairs;
    • profiles saved and reloaded before merging;
    • other list and scalar patterns.

17 of the new tests fail on the previous code.

Verification

  • ruff check . and mypy src/freshdata: clean.
  • pytest -m "not online and not large":
    • Python 3.12 / pandas 2.3.3: 6282 passed.
    • Python 3.9 / pandas 1.5.3: 6253 passed.
  • pandas 1.x can't clean Arrow string columns at all, even without a profile, so on 1.x the tests exercise only the drift gate for those dtypes.

The drift gate only treated object as text, so a profile learned on a
categorical or string column and replayed on the same values as object
(or the reverse) reported mild drift and dropped that column's value
maps and hints. A text categorical replayed on numeric data was also
not flagged.

Compare text-ness with _util.is_text_dtype on the frame's dtype and a
name-based equivalent for the stored learned dtype. object, string,
Arrow string and text categoricals are interchangeable; text vs
non-text is still reported as drift.
union_min_precision and error_on_conflict assumed every entry of the
embedded memory's value_patterns was a {raw: clean} mapping and raised
ValueError when memory held the semantic_repairs list.

Union list patterns with de-duplication. A semantic repair both sides
propose differently is dropped and recorded under union_min_precision
and raises ProfileMergeError under error_on_conflict; differing scalar
patterns are handled the same way. prefer_self/prefer_other still copy
the preferred side's memory.
The masking salt comes from the frame signature (row count, dtypes and
a head sample), so the same values can mask to different tokens in
different profiles and merged profiles cannot match each other's
masked entries.
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 9c99f349-441c-454f-a9be-31e4f3586b50


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit ffab2a2 into main Sep 15, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant