Skip to content

Commit 204a3f4

Browse files
fix(freshcore): fall back when native casts create columns the kernels mishandle (#422)
The adapter's input checks only see the input dtypes, but FreshCore's native fix_dtypes stage casts text columns before imputation and outlier detection run. Two casts reach known kernel gaps: - "yes"/"no" text with missing values becomes a boolean column, which the native imputer skips, so mode/auto imputation leaves it unfilled while pandas fills it. - numeric text holding "inf" becomes a float column holding inf, which the native fences do not exclude, so zscore/iqr flag nothing while pandas drops inf before fencing. After the native run, the adapter now checks the returned column dtypes for text columns cast to bool or float. When a cast column hits either gap, it records a fallback naming the column (step "impute" or "outliers") and reruns on pandas. Under fallback_policy="error" it raises FallbackError. The scan covers only cast columns and runs only when impute or outliers is set, so other frames stay native.
1 parent 757a24e commit 204a3f4

5 files changed

Lines changed: 317 additions & 0 deletions

File tree

‎docs/fallback-matrix.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -26,10 +26,12 @@ FreshCore also runs its own config and data checks in
2626
| detection-only dedup (`drop_duplicates=False`) with `duplicate_ratio_action="error"` | native | native | native | native (pandas with native modules that don't report `duplicates_detected`) | the escalation needs the duplicate-row count at the pandas dedup stage; FreshCore counts it natively, but modules built before that count existed would never raise |
2727
| global impute mean/median/mode | native | native | native | native | — |
2828
| impute `mode`/`auto` with missing values in a nullable `boolean` column | native | native | native | pandas | FreshCore v1 kernels do not impute boolean columns |
29+
| impute `mode`/`auto` when `fix_dtypes` casts a text column (e.g. `"yes"`/`"no"`) to boolean and it still has missing values | native | native | native | pandas | same kernel gap as above, but the column only becomes boolean inside the native run, so the adapter checks the native result (`fallback_step="impute"`) and reruns on pandas |
2930
| per-column `impute_strategy` | pandas | pandas | pandas | pandas | unimplemented natively (no fundamental blocker) |
3031
| `impute="missforest"` | pandas | pandas | pandas | pandas | scikit-learn model |
3132
| outliers `iqr` / `zscore` | native | native | native | native | — |
3233
| outliers when a float column holds `±inf` | native | native | native | pandas | FreshCore v1 fences don't exclude non-finite values, so they flag or clip nothing |
34+
| outliers when `fix_dtypes` casts a text column holding `"inf"`/`"-inf"` to float | native | native | native | pandas | same fence gap as above, but the infinity only appears inside the native run, so the adapter checks the native result (`fallback_step="outliers"`) and reruns on pandas |
3335
| `outlier_action="auto"` / model methods | pandas | pandas | pandas | pandas | data-dependent / model-based selection |
3436
| `fix_dtypes=True` (default) | pandas | pandas | pandas | partial | sampled heuristics on the pandas reference; FreshCore casts bool/numeric natively, defers datetimes |
3537
| `drop_constant_columns` | pandas | pandas | pandas | pandas | needs a data scan before planning (two-phase plan not built) |

‎docs/freshcore.md‎

Lines changed: 12 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -57,6 +57,18 @@ back for `impute="missforest"`, per-column `impute_strategy`, outlier handling
5757
on float columns holding `±inf`, and mode/auto imputation of nullable boolean
5858
columns with missing values.
5959

60+
The same two gaps can open inside a native run. `fix_dtypes` casts text columns
61+
before imputation and outlier handling, so a `"yes"`/`"no"` column with missing
62+
values becomes a boolean column the kernels do not impute. A numeric text
63+
column holding `"inf"` becomes a float column holding `±inf`, which leaves the
64+
outlier fences undefined. The input frame shows neither, so the adapter checks
65+
the returned column dtypes. When a cast column hits a gap, it reruns the frame
66+
on pandas and records a fallback event naming the column
67+
(`fallback_step="impute"` or `"outliers"`). Under `fallback_policy="error"` it
68+
raises `FallbackError` instead. `fd.plan()` checks only the input, so it
69+
cannot predict these fallbacks, and the discarded native run still costs time.
70+
Frames without such casts stay on the native path.
71+
6072
With `drop_duplicates=False` (the default), the native module counts full-row
6173
duplicates at the same stage as the pandas step: after string cleaning,
6274
empty-row removal and casts, and before imputation and outliers. It returns

‎src/freshdata/execution/backends/_freshcore.py‎

Lines changed: 57 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -105,6 +105,11 @@ def execute(
105105
"which this FreshCore native module does not report",
106106
)
107107

108+
cast_fallback = self._native_cast_reason(native, frame, config)
109+
if cast_fallback is not None:
110+
step, reason = cast_fallback
111+
return self._fallback(source, config, engine_config, step, reason)
112+
108113
cleaned, dtype_changes = self._frame_from_native(native, frame, config)
109114
report = self._report_from_native(frame, cleaned, native, started, config)
110115
for column, detail in dtype_changes:
@@ -306,6 +311,58 @@ def _boolean_columns_to_impute(frame: pd.DataFrame) -> list[str]:
306311
found.append(str(col))
307312
return found
308313

314+
def _native_cast_reason(
315+
self, native: dict[str, Any], frame: pd.DataFrame, config: CleanConfig
316+
) -> tuple[str, str] | None:
317+
"""Fallback ``(step, reason)`` when a native cast built a column the kernels mishandle.
318+
319+
``_unsupported_reason`` sees only the input dtypes, but the native
320+
``fix_dtypes`` stage casts text columns before imputation and outlier
321+
detection run. A text column cast to ``bool`` or ``float`` has the same
322+
gaps as a boolean or ±inf input column: missing booleans are never
323+
imputed, and ±inf leaves the outlier fences undefined. The returned
324+
column dtypes show which columns were cast, so only those are scanned,
325+
and only when the configured step would reach the gap.
326+
"""
327+
check_bools = config.impute in ("mode", "auto")
328+
check_inf = config.outliers is not None
329+
if not config.fix_dtypes or not (check_bools or check_inf):
330+
return None
331+
sources = {
332+
str(label): (label, frame.iloc[:, i])
333+
for i, label in enumerate(self._output_labels(frame, config))
334+
}
335+
bools: list[str] = []
336+
infinite: list[str] = []
337+
for column in native.get("columns", []):
338+
source = sources.get(column["name"])
339+
if source is None: # e.g. an outlier flag column added natively
340+
continue
341+
label, series = source
342+
dtype = column.get("dtype")
343+
values = column["values"]
344+
if check_bools and dtype == "bool" and not is_bool_dtype(series):
345+
n_missing = values.count(None)
346+
if 0 < n_missing < len(values):
347+
bools.append(str(label))
348+
elif check_inf and dtype == "float" and not is_numeric_dtype(series):
349+
floats = np.asarray(values, dtype="float64") # None -> NaN
350+
if np.isinf(floats).any():
351+
infinite.append(str(label))
352+
if bools:
353+
return "impute", (
354+
f"text column(s) {self._shown(bools)} were cast to boolean by FreshCore v1 "
355+
"and still hold missing values: FreshCore v1 does not impute boolean "
356+
"columns, so imputation requires the pandas reference path"
357+
)
358+
if infinite:
359+
return "outliers", (
360+
f"text column(s) {self._shown(infinite)} were cast to float64 by FreshCore v1 "
361+
"and hold ±inf: FreshCore v1 outlier fences do not exclude ±inf, so outlier "
362+
"handling requires the pandas reference path"
363+
)
364+
return None
365+
309366
@staticmethod
310367
def _has_unsupported_object_values(frame: pd.DataFrame) -> bool:
311368
for col in frame.columns:

‎tests/test_execution/test_freshcore_engine.py‎

Lines changed: 129 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,7 @@
77
from __future__ import annotations
88

99
import pandas as pd
10+
import pytest
1011

1112
import freshdata as fd
1213
from freshdata.execution import EngineConfig, EngineSelector
@@ -99,6 +100,134 @@ def execute_plan(payload):
99100
assert report.to_dict()["stage_timings"][0]["backend"] == "freshcore"
100101

101102

103+
class _CastingNative:
104+
"""Echoes the input but reports *cast* columns as a native fix_dtypes would."""
105+
106+
calls = 0
107+
cast: dict = {}
108+
109+
@classmethod
110+
def execute_plan(cls, payload):
111+
cls.calls += 1
112+
columns = [cls.cast.get(c["name"], c) for c in payload["columns"]]
113+
return {"columns": columns, "actions": []}
114+
115+
116+
def _casting(monkeypatch, **cast) -> type[_CastingNative]:
117+
native = type("CastingNative", (_CastingNative,), {"calls": 0, "cast": cast})
118+
monkeypatch.setattr(FreshCoreEngine, "_load_native", staticmethod(lambda: native))
119+
return native
120+
121+
122+
def _cast_cfg(**kwargs) -> fd.CleanConfig:
123+
return fd.CleanConfig(strategy="conservative", verbose=False, **kwargs)
124+
125+
126+
_YES_NO = pd.DataFrame({"flag": ["yes", "no", None, "yes"], "k": [1.0, 2, 3, 4]})
127+
_BOOL_CAST = {"name": "flag", "dtype": "bool", "values": [True, False, None, True]}
128+
_INF_TEXT = pd.DataFrame({"x": ["1", "2", "3", "100", "inf"]})
129+
_INF_CAST = {"name": "x", "dtype": "float", "values": [1.0, 2.0, 3.0, 100.0, float("inf")]}
130+
131+
132+
@pytest.mark.parametrize("impute", ["mode", "auto"])
133+
def test_native_bool_cast_with_missing_values_falls_back_for_imputation(monkeypatch, impute):
134+
native = _casting(monkeypatch, flag=_BOOL_CAST)
135+
out, report = fd.clean(
136+
_YES_NO.copy(), config=_cast_cfg(impute=impute), engine="freshcore", return_report=True
137+
)
138+
139+
assert native.calls == 1
140+
assert report.backend == "pandas"
141+
[event] = report.fallback_events
142+
assert event["fallback_step"] == "impute"
143+
assert "'flag'" in event["fallback_reason"]
144+
assert "cast to boolean" in event["fallback_reason"]
145+
assert out["flag"].tolist() == [True, False, True, True]
146+
147+
148+
@pytest.mark.parametrize("method", ["zscore", "iqr"])
149+
def test_native_float_cast_holding_inf_falls_back_for_outliers(monkeypatch, method):
150+
native = _casting(monkeypatch, x=_INF_CAST)
151+
_, report = fd.clean(
152+
_INF_TEXT.copy(),
153+
config=_cast_cfg(outliers="flag", outlier_method=method),
154+
engine="freshcore",
155+
return_report=True,
156+
)
157+
158+
assert native.calls == 1
159+
assert report.backend == "pandas"
160+
[event] = report.fallback_events
161+
assert event["fallback_step"] == "outliers"
162+
assert "'x'" in event["fallback_reason"]
163+
assert "±inf" in event["fallback_reason"]
164+
165+
166+
@pytest.mark.parametrize(
167+
("frame", "cast", "options"),
168+
[
169+
pytest.param(_YES_NO, {"flag": _BOOL_CAST}, {"impute": "mode"}, id="bool"),
170+
pytest.param(_INF_TEXT, {"x": _INF_CAST}, {"outliers": "flag"}, id="inf"),
171+
],
172+
)
173+
def test_native_cast_fallback_honours_error_policy(monkeypatch, frame, cast, options):
174+
_casting(monkeypatch, **cast)
175+
with pytest.raises(fd.FallbackError):
176+
fd.clean(
177+
frame.copy(),
178+
config=_cast_cfg(**options),
179+
engine="freshcore",
180+
fallback_policy="error",
181+
)
182+
183+
184+
@pytest.mark.parametrize(
185+
("frame", "cast", "options"),
186+
[
187+
pytest.param(
188+
_YES_NO, {"flag": _BOOL_CAST}, {"outliers": "flag"}, id="bool-cast-without-impute"
189+
),
190+
pytest.param(
191+
_YES_NO, {"flag": _BOOL_CAST}, {"impute": "median"}, id="bool-cast-median-skips"
192+
),
193+
pytest.param(
194+
_YES_NO,
195+
{"flag": {**_BOOL_CAST, "values": [True, False, False, True]}},
196+
{"impute": "mode"},
197+
id="bool-cast-without-missing",
198+
),
199+
pytest.param(
200+
_INF_TEXT, {"x": _INF_CAST}, {"impute": "median"}, id="inf-cast-without-outliers"
201+
),
202+
pytest.param(
203+
_INF_TEXT,
204+
{"x": {**_INF_CAST, "values": [1.0, 2.0, 3.0, 100.0, None]}},
205+
{"outliers": "flag"},
206+
id="finite-float-cast",
207+
),
208+
pytest.param(
209+
pd.DataFrame({"x": ["a", "b", None], "k": [1.0, 2, 3]}),
210+
{},
211+
{"impute": "mode", "outliers": "flag"},
212+
id="no-cast",
213+
),
214+
],
215+
)
216+
def test_frames_without_mishandled_casts_stay_native(monkeypatch, frame, cast, options):
217+
native = _casting(monkeypatch, **cast)
218+
_, report = fd.clean(
219+
frame.copy(),
220+
config=_cast_cfg(**options),
221+
engine="freshcore",
222+
fallback_policy="error",
223+
return_report=True,
224+
)
225+
226+
assert native.calls == 1
227+
assert report.backend == "freshcore"
228+
assert report.fallback_events == []
229+
230+
102231
def test_string_case_available_on_reference_pipeline():
103232
df = pd.DataFrame({"name": ["Alice", "BOB"]})
104233
out, report = fd.clean(df, config=_cfg(string_case="lower"), return_report=True)

‎tests/test_execution/test_freshcore_native_parity.py‎

Lines changed: 117 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -256,3 +256,120 @@ def test_detection_counts_before_imputation_like_pandas(native):
256256
df = pd.DataFrame({"a": [1.0, 2.0, None, 2.0], "b": ["x", "y", "y", "z"]})
257257
expected, report = _clean_both(native, df, impute="mean", **REPRO)
258258
assert _detections(report) == _detections(expected) == []
259+
260+
261+
# -- native casts that build columns the kernels mishandle -------------------
262+
#
263+
# ``fix_dtypes`` runs natively, so a text column can turn into a boolean column
264+
# with missing values (never imputed natively) or a float column holding ±inf
265+
# (which leaves the native outlier fences undefined). The input frame shows
266+
# neither, so the adapter checks the native result and falls back.
267+
268+
269+
def _yes_no_frame(missing: bool = True) -> pd.DataFrame:
270+
flags = ["yes", "no", None if missing else "no", "yes", "Yes"]
271+
return pd.DataFrame({"flag": flags, "k": [1.0, 2, 3, 4, 5]})
272+
273+
274+
def _inf_text_frame(values: list[str] | None = None) -> pd.DataFrame:
275+
if values is None:
276+
values = ["1", "2", "3", "4", "5", "6", "7", "8", "9", "100", "inf"]
277+
return pd.DataFrame({"x": values})
278+
279+
280+
INF_ZSCORE = {"outliers": "flag", "outlier_method": "zscore", "outlier_factor": 2.0}
281+
CAST_REPROS = [
282+
pytest.param(_yes_no_frame(), {"impute": "mode"}, "impute", "'flag'", id="bool-mode"),
283+
pytest.param(_yes_no_frame(), {"impute": "auto"}, "impute", "'flag'", id="bool-auto"),
284+
pytest.param(_inf_text_frame(), INF_ZSCORE, "outliers", "'x'", id="inf-zscore"),
285+
pytest.param(
286+
_inf_text_frame(["1", "2", "3", "4", "inf", "inf", "inf", "inf"]),
287+
{"outliers": "flag", "outlier_method": "iqr"},
288+
"outliers",
289+
"'x'",
290+
id="inf-iqr",
291+
),
292+
pytest.param(
293+
_inf_text_frame(["1", "2", "3", "4", "5", "6", "7", "8", "9", "100", "-inf"]),
294+
INF_ZSCORE,
295+
"outliers",
296+
"'x'",
297+
id="negative-inf-zscore",
298+
),
299+
]
300+
301+
302+
def test_native_casts_diverge_without_the_fallback():
303+
"""Pins the kernel behaviour the adapter guards against."""
304+
booleans = _native_result(_yes_no_frame(), impute="mode")
305+
[flag] = [c for c in booleans["columns"] if c["name"] == "flag"]
306+
assert flag["dtype"] == "bool"
307+
assert None in flag["values"]
308+
309+
infinite = _native_result(_inf_text_frame(), **INF_ZSCORE)
310+
assert [c["name"] for c in infinite["columns"]] == ["x"] # no flag column
311+
assert infinite["outliers_handled"] == 0
312+
313+
314+
@pytest.mark.parametrize(("df", "options", "step", "column"), CAST_REPROS)
315+
def test_native_cast_repros_fall_back_to_the_pandas_result(native, df, options, step, column):
316+
expected, _ = fd.clean(df.copy(), engine="pandas", return_report=True, **KW, **options)
317+
out, report = fd.clean(df.copy(), engine="freshcore", return_report=True, **KW, **options)
318+
319+
assert native.calls == 1
320+
assert report.backend == "pandas"
321+
[event] = report.fallback_events
322+
assert event["fallback_step"] == step
323+
assert column in event["fallback_reason"]
324+
pd.testing.assert_frame_equal(pd.DataFrame(out), pd.DataFrame(expected))
325+
326+
327+
def test_native_cast_repro_values_match_pandas(native):
328+
out = fd.clean(_yes_no_frame(), engine="freshcore", impute="mode", **KW)
329+
assert out["flag"].tolist() == [True, False, True, True, True]
330+
331+
flagged = fd.clean(_inf_text_frame(), engine="freshcore", **INF_ZSCORE, **KW)
332+
assert flagged["x_outlier"].tolist() == [False] * 9 + [True, True]
333+
334+
335+
@pytest.mark.parametrize(("df", "options", "step", "column"), CAST_REPROS)
336+
def test_native_cast_repros_raise_under_error_policy(native, df, options, step, column):
337+
with pytest.raises(fd.FallbackError, match=f"step {step!r}"):
338+
fd.clean(df.copy(), engine="freshcore", fallback_policy="error", **KW, **options)
339+
assert native.calls == 1
340+
341+
342+
@pytest.mark.parametrize(
343+
("df", "options"),
344+
[
345+
pytest.param(_yes_no_frame(missing=False), {"impute": "mode"}, id="bool-cast-no-missing"),
346+
pytest.param(_yes_no_frame(), {"outliers": "flag"}, id="bool-cast-without-impute"),
347+
pytest.param(_inf_text_frame(), {"impute": "median"}, id="inf-cast-without-outliers"),
348+
pytest.param(
349+
_inf_text_frame(["1", "2", "3", "4", "5", "6", "7", "8", "9", "100", None]),
350+
{**INF_ZSCORE, "impute": "median"},
351+
id="finite-float-cast",
352+
),
353+
pytest.param(
354+
pd.DataFrame({"s": ["a", "b", None, "a"], "v": [1.0, 2.0, 3.0, 40.0]}),
355+
{"impute": "mode", "outliers": "flag"},
356+
id="no-cast",
357+
),
358+
],
359+
)
360+
def test_frames_without_mishandled_casts_stay_native(native, df, options):
361+
expected, _ = fd.clean(df.copy(), engine="pandas", return_report=True, **KW, **options)
362+
out, report = fd.clean(
363+
df.copy(),
364+
engine="freshcore",
365+
return_report=True,
366+
fallback_policy="error",
367+
**KW,
368+
**options,
369+
)
370+
371+
assert native.calls == 1
372+
assert report.backend == "freshcore"
373+
assert report.fallback_events == []
374+
# Values match; dtypes may not (e.g. native boolean vs pandas bool).
375+
pd.testing.assert_frame_equal(pd.DataFrame(out), pd.DataFrame(expected), check_dtype=False)

0 commit comments

Comments
 (0)