Skip to content

Commit 9368ea4

Browse files
Merge pull request #151 from FreshCode-Org/fix/vectorize-scinot-guard-jwd
perf(dtypes): fix T5 runtime regression from the scientific-notation guard
2 parents 497c2f6 + f8ecd1a commit 9368ea4

2 files changed

Lines changed: 27 additions & 15 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -96,6 +96,15 @@ adheres to [Semantic Versioning](https://semver.org/).
9696
documentation site (`docs/superpowers/`).
9797

9898
### Fixed
99+
- **Default-path slowdown from the scientific-notation segfault guard**
100+
(release blocker, nightly issue #147): the guard that masks huge-exponent
101+
tokens (`"1e999"`) before `pd.to_numeric` — protection against a pandas
102+
2.3.x segfault — screened text columns cell by cell through a Python
103+
predicate, roughly doubling `fix_dtypes` time on 50k-row frames in CI.
104+
Each column is now screened with a single C-level joined-blob regex scan
105+
and the per-cell predicate runs only on columns that screen positive.
106+
Masking semantics are unchanged; the CleanBench T5 runtime gate is back
107+
within its ±20 % baseline envelope.
99108
- `dir(freshdata)` no longer lists `Action` twice: the privacy-policy engine's
100109
`Action` enum was listed in the lazy enterprise exports but was unreachable
101110
there — `fd.Action` is (and remains) the audit action from

‎src/freshdata/steps/dtypes.py‎

Lines changed: 18 additions & 15 deletions
Original file line numberDiff line numberDiff line change
@@ -42,6 +42,12 @@
4242
r"^[+-]?(?:\d+(?:\.\d*)?|\.\d+)[eE]([+-]?\d+)$"
4343
)
4444
_MAX_SAFE_EXPONENT = 308
45+
# Column-level screen for the guard in _to_numeric_or_none: any string that
46+
# pandas could parse as scientific notation with an exponent past the float
47+
# range must contain an exponent marker followed by at least three digits
48+
# (309 is the smallest unsafe magnitude). A false positive merely routes the
49+
# column to the exact per-cell check.
50+
_RISKY_EXPONENT = re.compile(r"[eE][+-]?\d{3,}")
4551

4652

4753
def _number_format(
@@ -153,22 +159,19 @@ def _to_numeric_or_none(values: pd.Series) -> pd.Series | None:
153159
if pd.api.types.is_object_dtype(values.dtype) or pd.api.types.is_string_dtype(
154160
values.dtype
155161
):
156-
# Only strings containing an exponent marker can match the unsafe
157-
# pattern, so find candidates with one vectorized pass and run the
158-
# per-value regex on that (normally empty) subset only. ``.str``
159-
# refuses object columns that contain no strings at all — such a
160-
# column has no unsafe tokens either, so treat it as candidate-free.
162+
# Screen the whole column as one joined blob first: a single C-level
163+
# join plus one regex scan, no per-cell Python work in the common
164+
# (safe) case. The join raises TypeError when non-string, non-missing
165+
# objects are present — treat such columns as risky and let the exact
166+
# per-cell predicate decide.
161167
try:
162-
candidates = values.str.contains("e", case=False, regex=False, na=False)
163-
except (AttributeError, TypeError):
164-
candidates = None
165-
if candidates is not None:
166-
if candidates.dtype != bool:
167-
candidates = candidates.fillna(False).astype(bool)
168-
if bool(candidates.any()):
169-
unsafe = values[candidates].map(_has_unsafe_scientific_exponent)
170-
if bool(unsafe.any()):
171-
values = values.mask(unsafe.reindex(values.index, fill_value=False))
168+
blob = "\x1f".join(values.dropna().to_numpy())
169+
except TypeError:
170+
blob = None
171+
if blob is None or _RISKY_EXPONENT.search(blob) is not None:
172+
unsafe = values.map(_has_unsafe_scientific_exponent)
173+
if bool(unsafe.any()):
174+
values = values.mask(unsafe)
172175
try:
173176
return pd.to_numeric(values, errors="coerce")
174177
except (TypeError, ValueError):

0 commit comments

Comments
 (0)