Skip to content

Commit d98c970

Browse files
docs: correct three claims that do not match the code
Phase 32 truth-up. Each was verified against the code before being changed, and one of the three turned out to be less wrong than previously recorded. StreamingCleanConfig.window_size was documented as sizing "rolling statistics and the rolling trust score". It does neither. Its only functional use in the whole src tree is bounding the cross-batch duplicate window, and the rolling trust score is sized by a different field, rolling_trust_window. A user reading help(StreamingCleanConfig) would raise window_size expecting more history to feed trust scoring, and nothing would change. The prose docs in docs/streaming.md were already correct; the API docstring was the wrong one. docs/repair-plans.md described the drift fingerprint as "row count, column names+dtypes, content sample". That is accurate but incomplete in a way that matters: the sample is the first 512 rows, so a change after row 512 that preserves the row count, names and dtypes is not detected. Demonstrated: on a 1000-row frame the same edit raises PlanDriftError at row 10 and passes silently at row 900. The fingerprint is deliberately cheap and this is a design choice, not a bug -- but the head restriction belongs where the guarantee is described. A new test pins the boundary exactly at row 512 so the documented claim is backed by execution, and so a future change to the sampling strategy has to come past a failing test. The README's "Native Polars DataFrames" section showed fd.clean(pl_df) with defaults and called it native. docs/fallback-matrix.md opens by calling this "the single most important row": with default options every native engine delegates the whole pipeline to pandas. You get a Polars frame back, but not native Polars execution. That sentence now appears beside the example. Documentation and one new test only; no behaviour change.
1 parent 4b75cd7 commit d98c970

5 files changed

Lines changed: 112 additions & 5 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 15 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,21 @@ adheres to [Semantic Versioning](https://semver.org/).
66

77
## [Unreleased]
88

9+
### Documentation
10+
11+
- `StreamingCleanConfig.window_size` was documented as sizing "rolling
12+
statistics and the rolling trust score". It does neither: its only effect is
13+
to bound the cross-batch duplicate window, and the rolling trust score is
14+
sized by `rolling_trust_window`. Docstring corrected; no behaviour change.
15+
- `docs/repair-plans.md` now states that the `FrameSignature` content sample is
16+
the first 512 rows, so a change beyond row 512 that preserves row count,
17+
column names and dtypes is not detected by drift refusal.
18+
- The README's "Native Polars DataFrames" section now states what
19+
`docs/fallback-matrix.md` already did: with default options every native
20+
engine delegates the whole pipeline to pandas, and the fully native path is
21+
`strategy="conservative"` with `fix_dtypes=False`.
22+
23+
924
### Fixed
1025
- A domain regex rule no longer fails a valid code because its column is
1126
`float64`. A numeric code column with one blank cell loads from CSV as

‎README.md‎

Lines changed: 8 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -224,6 +224,14 @@ cleaned = fd.clean(df)
224224

225225
### 2. Native Polars DataFrames
226226
Pass a Polars DataFrame, get a Polars DataFrame back with zero pandas boilerplate:
227+
228+
> **With default options the work still runs on pandas.** The default
229+
> `strategy="balanced"` runs the accuracy-first decision engine, which is
230+
> evaluated by the pandas backend, so every native engine delegates the whole
231+
> pipeline to pandas and records it on `report.fallback_events`. You get a
232+
> Polars frame back, but not native Polars execution. The fully native path is
233+
> `strategy="conservative"` with `fix_dtypes=False`. See
234+
> [docs/fallback-matrix.md](docs/fallback-matrix.md).
227235
```python
228236
import polars as pl
229237
import freshdata as fd

‎docs/repair-plans.md‎

Lines changed: 6 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -86,8 +86,12 @@ everything non-trivial stays `pending`.
8686
## Drift refusal
8787

8888
A plan remembers the frame it was built for (`FrameSignature`: row count,
89-
column names+dtypes, content sample). Applying it to different data refuses
90-
by default:
89+
column names+dtypes, content sample). The content sample is the **first 512
90+
rows** (`_SIGNATURE_SAMPLE_ROWS`), not the whole frame: a change beyond row
91+
512 that leaves the row count, column names and dtypes intact is **not**
92+
detected. The fingerprint is deliberately cheap; use it as a guard against
93+
applying a plan to the wrong data, not as a proof the data is unchanged.
94+
Applying a plan to different data refuses by default:
9195

9296
```python
9397
fd.apply_plan(other_df, rp) # raises fd.PlanDriftError

‎src/freshdata/streaming/_config.py‎

Lines changed: 5 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -19,9 +19,11 @@ class StreamingCleanConfig:
1919
Parameters
2020
----------
2121
window_size:
22-
Size of the recent-window used for rolling statistics and the rolling
23-
trust score. The *caller* controls how big each batch is; this only
24-
bounds how much recent history influences "recent-window" reporting.
22+
Bound on the cross-batch duplicate window: the number of most recently
23+
seen distinct rows retained when ``global_duplicates`` is enabled. A
24+
duplicate older than the window is not detected. This is its only
25+
effect -- it does **not** size rolling statistics, and the rolling
26+
trust score is sized by ``rolling_trust_window`` instead.
2527
warmup_batches:
2628
Number of leading batches during which the cleaner only repairs
2729
representation and *collects* statistics — it defers statistical

‎tests/test_plan_drift_sampling.py‎

Lines changed: 78 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,78 @@
1+
"""The frame fingerprint behind drift refusal samples the HEAD, not the frame.
2+
3+
``docs/repair-plans.md`` describes ``FrameSignature`` as "row count, column
4+
names+dtypes, content sample". That wording is accurate but incomplete in a way
5+
that matters: the sample is the first ``_SIGNATURE_SAMPLE_ROWS`` (512) rows, so
6+
a change *after* row 512 that preserves the row count, the column names and the
7+
dtypes does not trip ``PlanDriftError``.
8+
9+
That is a deliberate design choice -- the fingerprint is documented as "cheap"
10+
and is a guard against applying a plan to the *wrong data*, not a proof the
11+
data is unchanged. These tests pin the boundary so the documented claim is
12+
backed by execution rather than prose, and so a future change to the sampling
13+
strategy has to come past a failing test.
14+
"""
15+
16+
from __future__ import annotations
17+
18+
import pandas as pd
19+
import pytest
20+
21+
import freshdata as fd
22+
from freshdata.repairplan import _SIGNATURE_SAMPLE_ROWS, compute_frame_signature
23+
24+
25+
def _frame(n: int) -> pd.DataFrame:
26+
# A tie-free majority so the plan carries a stable semantic action.
27+
lower = max(1, n // 4)
28+
return pd.DataFrame({"country": ["USA"] * (n - lower) + ["usa"] * lower})
29+
30+
31+
def _plan(df: pd.DataFrame):
32+
return fd.suggest_plan(df, semantic_mode="review")
33+
34+
35+
def test_the_documented_sample_size_is_the_one_the_code_uses():
36+
assert _SIGNATURE_SAMPLE_ROWS == 512
37+
38+
39+
def test_a_change_inside_the_head_sample_is_refused():
40+
df = _frame(1000)
41+
plan = _plan(df)
42+
drifted = df.copy()
43+
drifted.loc[10, "country"] = "COMPLETELY-DIFFERENT"
44+
with pytest.raises(fd.PlanDriftError):
45+
fd.apply_plan(drifted, plan)
46+
47+
48+
def test_a_change_beyond_the_head_sample_is_not_detected():
49+
"""Documented limitation, pinned deliberately -- not an endorsement."""
50+
df = _frame(1000)
51+
plan = _plan(df)
52+
drifted = df.copy()
53+
drifted.loc[900, "country"] = "COMPLETELY-DIFFERENT"
54+
fd.apply_plan(drifted, plan) # no PlanDriftError
55+
56+
57+
def test_the_boundary_sits_exactly_at_the_sample_size():
58+
df = _frame(_SIGNATURE_SAMPLE_ROWS + 10)
59+
base = compute_frame_signature(df)
60+
61+
last_seen = df.copy()
62+
last_seen.loc[_SIGNATURE_SAMPLE_ROWS - 1, "country"] = "CHANGED"
63+
assert compute_frame_signature(last_seen).sample_hash != base.sample_hash
64+
65+
first_unseen = df.copy()
66+
first_unseen.loc[_SIGNATURE_SAMPLE_ROWS, "country"] = "CHANGED"
67+
assert compute_frame_signature(first_unseen).sample_hash == base.sample_hash
68+
69+
70+
def test_row_count_and_column_changes_are_still_caught_beyond_the_sample():
71+
"""The sample is only one of three components; the other two still apply."""
72+
df = _frame(_SIGNATURE_SAMPLE_ROWS + 10)
73+
base = compute_frame_signature(df)
74+
75+
assert compute_frame_signature(df.iloc[:-1]).n_rows != base.n_rows
76+
77+
renamed = df.rename(columns={"country": "nation"})
78+
assert compute_frame_signature(renamed).columns_hash != base.columns_hash

0 commit comments

Comments
 (0)