Skip to content

Commit f9f97c3

Browse files
fix(dtypes): dayfirst=True no longer corrupts ISO-8601 dates
fd.clean(df, dayfirst=True) on ["2021-01-05", "2021-02-11"] silently returned 2021-05-01 and 2021-11-02 -- month and day swapped, with no warning, no report entry and no coercion record. dayfirst exists to resolve ambiguous short dates like 05/12/2021. An ISO-8601 date is YYYY-MM-DD by definition and has no ambiguity to resolve, so this was not a defensible reading of the input. Cause: pandas 2 infers one format for a whole column from its first value, and under dayfirst=True reads an ISO date as %Y-%d-%m. _parse_datetime now passes format="mixed" in that case, so each value is read by its own shape. pandas 1.x infers per value and was never affected -- verified directly on 1.5.3 -- and has no format="mixed", hence the version guard. What made it easy to miss is that the corruption was data-dependent. It was silent only while every day was <= 12: a day >= 13 made the guessed format fail on that value, dropped the parse share below datetime_threshold, and triggered the mixed-format retry that produced the correct reading. The same column therefore read correctly or incorrectly depending on values it happened to contain, or on an unrelated threshold. A user validating on a sample containing a day >= 13 would see correct output and ship. Only the top-level dayfirst=True kwarg was affected. dayfirst="auto" and both semantic_context routes were already correct, which is why this was never caught. Verified the tests fail without the fix rather than assuming: reverting src gives 3 failed / 5 passed. With it, 8 pass on py3.12/pandas 2.3.3 and on py3.9/pandas 1.5.3. dayfirst still resolves genuinely ambiguous slash dates in both directions.
1 parent e1520d7 commit f9f97c3

3 files changed

Lines changed: 119 additions & 1 deletion

File tree

‎CHANGELOG.md‎

Lines changed: 28 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,34 @@ adheres to [Semantic Versioning](https://semver.org/).
66

77
## [Unreleased]
88

9+
### Fixed
10+
11+
- `dayfirst=True` no longer reinterprets unambiguous ISO-8601 dates.
12+
`fd.clean(df, dayfirst=True)` on `["2021-01-05", "2021-02-11"]` silently
13+
returned `2021-05-01`, `2021-11-02` — month and day swapped, with no warning
14+
and no coercion record. `dayfirst` resolves *ambiguous* short dates such as
15+
`05/12/2021`; an ISO-8601 date is `YYYY-MM-DD` by definition and has no
16+
ambiguity to resolve. Cause: pandas 2 infers one format for a whole column
17+
from its first value and, under `dayfirst=True`, reads an ISO date as
18+
`%Y-%d-%m`; `_parse_datetime` now passes `format="mixed"` in that case so
19+
each value is read by its own shape. pandas 1.x infers per value and was
20+
never affected, and has no `format="mixed"`, so the change is guarded on the
21+
pandas major version.
22+
23+
The corruption was data-dependent, which is what made it easy to miss: it was
24+
silent only while every day was `<= 12`, because a day `>= 13` made the
25+
guessed format fail, dropped the parse share below `datetime_threshold`, and
26+
triggered the mixed-format retry that produced the correct reading. The same
27+
column therefore read correctly or incorrectly depending on values it
28+
happened to contain, or on an unrelated threshold.
29+
30+
**Compatibility impact:** output changes for ISO-8601 columns cleaned with
31+
`dayfirst=True` on pandas 2 — from a wrong reading to the correct one.
32+
`dayfirst` behaviour on genuinely ambiguous slash dates is unchanged, as are
33+
the `dayfirst="auto"` and `semantic_context` routes, which were never
34+
affected.
35+
36+
937
### Documentation
1038

1139
- `StreamingCleanConfig.window_size` was documented as sizing "rolling

‎src/freshdata/steps/dtypes.py‎

Lines changed: 10 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -381,7 +381,16 @@ def _parse_datetime(
381381
s: pd.Series, mixed_formats: bool, dayfirst: bool = False
382382
) -> pd.Series | None:
383383
kwargs: dict = {"errors": "coerce", "dayfirst": dayfirst}
384-
if mixed_formats:
384+
if mixed_formats or (dayfirst and PANDAS_MAJOR >= 2):
385+
# ``dayfirst`` exists to resolve *ambiguous* short dates like
386+
# ``05/12/2021``. An ISO-8601 date is not ambiguous. But pandas 2's
387+
# format inference infers a single format for the whole column from
388+
# the first value, and with ``dayfirst=True`` it reads ``2021-01-05``
389+
# as ``%Y-%d-%m`` -- silently returning 2021-05-01. ``format="mixed"``
390+
# parses each value by its own apparent format, which fixes ISO input
391+
# while leaving genuinely ambiguous slash dates to ``dayfirst``.
392+
# pandas 1.x infers per value already and is unaffected; it also has
393+
# no ``format="mixed"``, hence the version guard.
385394
kwargs["format"] = "mixed"
386395
with warnings.catch_warnings():
387396
warnings.simplefilter("ignore") # format-inference chatter; report covers it

‎tests/test_dayfirst_iso_dates.py‎

Lines changed: 81 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,81 @@
1+
"""``dayfirst=True`` must not reinterpret unambiguous ISO-8601 dates.
2+
3+
``dayfirst`` exists to resolve *ambiguous* short dates such as ``05/12/2021``.
4+
An ISO-8601 date is ``YYYY-MM-DD`` by definition, so there is nothing for it to
5+
resolve. Before this fix ``fd.clean(df, dayfirst=True)`` silently returned
6+
``2021-05-01`` for the input ``2021-01-05``.
7+
8+
The cause was pandas 2's format inference, which picks ONE format for a whole
9+
column from its first value and, with ``dayfirst=True``, reads an ISO date as
10+
``%Y-%d-%m``. pandas 1.x infers per value and was never affected.
11+
12+
The corruption was data-dependent, which is what made it dangerous: it is
13+
silent only while every day is <= 12. A day >= 13 makes the guessed format fail
14+
on that value, dropping the parse share below ``datetime_threshold`` so the
15+
mixed-format retry replaces the column with the correct reading. The same
16+
column therefore read correctly or incorrectly depending on values it happened
17+
to contain, or on an unrelated threshold.
18+
"""
19+
20+
from __future__ import annotations
21+
22+
import pandas as pd
23+
import pytest
24+
25+
import freshdata as fd
26+
27+
ISO = ["2021-01-05", "2021-02-11", "2021-03-09"]
28+
29+
30+
def _dates(frame, column="v"):
31+
return [str(v.date()) for v in frame[column]]
32+
33+
34+
def test_iso_dates_are_not_reinterpreted_when_dayfirst_is_true():
35+
"""The regression. Before the fix this returned 2021-05-01 and friends."""
36+
out = fd.clean(pd.DataFrame({"v": ISO}), dayfirst=True, verbose=False)
37+
assert _dates(out) == ISO
38+
39+
40+
@pytest.mark.parametrize("dayfirst", [True, False, "auto", None])
41+
def test_iso_dates_read_the_same_whatever_dayfirst_says(dayfirst):
42+
"""``dayfirst`` is not a question ISO-8601 input can answer differently."""
43+
kwargs = {} if dayfirst is None else {"dayfirst": dayfirst}
44+
out = fd.clean(pd.DataFrame({"v": ISO}), verbose=False, **kwargs)
45+
assert _dates(out) == ISO
46+
47+
48+
def test_dayfirst_still_does_its_actual_job_on_ambiguous_slash_dates():
49+
"""The fix must not disarm ``dayfirst`` where it is genuinely needed."""
50+
ambiguous = ["05/12/2021", "06/11/2021"]
51+
day_first = fd.clean(pd.DataFrame({"v": ambiguous}), dayfirst=True, verbose=False)
52+
assert _dates(day_first) == ["2021-12-05", "2021-11-06"]
53+
54+
month_first = fd.clean(pd.DataFrame({"v": ambiguous}), dayfirst=False, verbose=False)
55+
assert _dates(month_first) == ["2021-05-12", "2021-06-11"]
56+
57+
58+
def test_the_reading_no_longer_depends_on_whether_a_day_exceeds_twelve():
59+
"""The data-dependence that made the corruption hard to notice.
60+
61+
Nineteen consecutive January dates. Before the fix these read correctly at
62+
the default threshold but became 2021-01-01, 2021-02-01, 2021-03-01 ... at
63+
``datetime_threshold=0.5``, losing seven values to ``NaT``.
64+
"""
65+
dates = [f"2021-01-{d:02d}" for d in range(1, 20)]
66+
67+
default = fd.clean(pd.DataFrame({"v": dates}), dayfirst=True, verbose=False)
68+
lowered = fd.clean(
69+
pd.DataFrame({"v": dates}), dayfirst=True, datetime_threshold=0.5, verbose=False
70+
)
71+
72+
assert _dates(default) == dates
73+
assert _dates(lowered) == dates
74+
assert not default["v"].isna().any()
75+
assert not lowered["v"].isna().any()
76+
77+
78+
def test_a_column_mixing_iso_and_ambiguous_forms_reads_each_by_its_own_shape():
79+
mixed = ["2021-01-05", "06/11/2021", "2021-03-09"]
80+
out = fd.clean(pd.DataFrame({"v": mixed}), dayfirst=True, verbose=False)
81+
assert _dates(out) == ["2021-01-05", "2021-11-06", "2021-03-09"]

0 commit comments

Comments
 (0)