Skip to content

fix(freshcore): fall back when native casts create columns the kernels mishandle - #422

Merged
kevincostner17 merged 1 commit into
mainfrom
fix/freshcore-fallback-after-casts
Sep 15, 2026
Merged

kevincostner17 merged 1 commit into
mainfrom
fix/freshcore-fallback-after-casts

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

FreshCore could return results that differ from pandas when its own fix_dtypes stage turned a text column into a column its kernels mishandle. The adapter now detects these casts in the native result and reruns the run on the pandas reference path, recording a fallback event that names the column.

Root cause

FreshCoreEngine._unsupported_reason checks only the input frame's dtypes. The native module casts text columns (casts.rs) before imputation and outlier detection run (python.rs), which opens two kernel gaps:

  • Booleans. try_bool turns "yes"/"no" text into a Bool column. missing.rs skips Bool columns, so under impute="mode"/"auto" their missing values stay unfilled. pandas fills them.
  • Infinity. parse_numeric accepts "inf". outliers.rs filters only NaN, so zscore/iqr fences on that column are undefined and nothing is flagged. pandas drops ±inf before computing fences.

Behaviour change

  • When the adapter falls back. With engine="freshcore", it reruns on pandas and records a fallback event naming the column in two cases:
    • fix_dtypes casts a text column to boolean, the column still has missing values, and impute is "mode" or "auto" (fallback_step="impute").
    • fix_dtypes casts a text column holding ±inf to float and outliers are enabled (fallback_step="outliers").
  • Fallback policy. fallback_policy="warn" warns; "error" raises FallbackError.
  • Scope of the check. It runs only when imputation or outliers are configured, and scans only the columns native cast. Other frames stay native. fd.plan() checks only the input, so it can't predict these fallbacks; docs/freshcore.md says so.
  • Docs. docs/fallback-matrix.md and docs/freshcore.md list the new fallbacks.

Default-output changes

None. The default strategy="balanced" already routes every engine="freshcore" run to pandas, and the defaults impute=None / outliers=None skip the new check. New fallback_events entries appear only when all of these hold:

  • engine="freshcore" with strategy="conservative"
  • impute="mode"/"auto", or outliers enabled
  • a cast column hits one of the gaps

There are no new backend_differences entries.

Tests

  • tests/test_execution/test_freshcore_engine.py tests the fallback decision with a fake module, so the native build isn't needed.
    • Both cases fall back, with the step and column in the reason, and "error" raises.
    • These stay native: boolean casts without impute or with median, casts without missing values, inf without outliers, finite float casts, and uncast frames.
  • tests/test_execution/test_freshcore_native_parity.py is skipped unless freshdata_freshcore is built.
    • It pins the current kernel behaviour.
    • The yes/no and inf repros (mode, auto, zscore, iqr, -inf) return the pandas frame via fallback, and "error" raises.
    • Unaffected frames stay native, with values matching pandas.

Verification

  • ruff check . and mypy src/freshdata: clean.
  • pytest -m "not online and not large": Python 3.12, 6269 passed; Python 3.9, 6240 passed.
  • With the native module built (maturin develop), pytest tests/test_execution -k freshcore gives 127 passed, 1 skipped.
  • Before and after on the native build:
    • yes/no + impute="mode": [T,F,<NA>,T,T] → [T,F,T,T,T], same as pandas.
    • "inf" + zscore/iqr: nothing flagged → flags match pandas.

…s mishandle

The adapter's input checks only see the input dtypes, but FreshCore's
native fix_dtypes stage casts text columns before imputation and outlier
detection run. Two casts reach known kernel gaps:

- "yes"/"no" text with missing values becomes a boolean column, which the
  native imputer skips, so mode/auto imputation leaves it unfilled while
  pandas fills it.
- numeric text holding "inf" becomes a float column holding inf, which the
  native fences do not exclude, so zscore/iqr flag nothing while pandas
  drops inf before fencing.

After the native run, the adapter now checks the returned column dtypes for
text columns cast to bool or float. When a cast column hits either gap, it
records a fallback naming the column (step "impute" or "outliers") and reruns
on pandas. Under fallback_policy="error" it raises FallbackError. The scan
covers only cast columns and runs only when impute or outliers is set, so
other frames stay native.
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 91928d1d-ee3e-4fd6-b833-58e2093da933


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit 204a3f4 into main Sep 15, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant