Skip to content

Commit 1ac83fc

Browse files
docs: changelog for the round-5 bug-hunt fixes (#392)
Add Unreleased Added, Changed and Fixed entries for merged PRs #351, #352, #353, #354, #356, #357, #358, #360, #361, #362, #363, #364, #365, #367, #368, #369, #370, #371, #372, #374, #375, #377, #378 and #379.
1 parent d837154 commit 1ac83fc

1 file changed

Lines changed: 199 additions & 0 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 199 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -16,6 +16,14 @@ adheres to [Semantic Versioning](https://semver.org/).
1616
tool orderings while keeping PyJanitor optional.
1717
- A dependency-optional Great Expectations recipe demonstrating the
1818
repair-then-validate workflow with an in-memory checkpoint.
19+
- `TimeSeriesCleanConfig.timestamp_unit` (`"s"`, `"ms"`, `"us"` or `"ns"`) sets
20+
the epoch unit for numeric timestamp and event-time columns in time-series
21+
streaming. Without it the unit is inferred from the values (#227).
22+
- `apply_review_decisions()` accepts a keyword-only `queue=` (the
23+
`ReviewQueueReport` the reviewer worked from), so decisions that carry only
24+
an `item_id` can be resolved to a record pair (#267).
25+
- The `FRESHDATA_MODEL_TIMEOUT` environment variable sets the network timeout,
26+
in seconds, for `fd.models.pull` downloads (default 60) (#341).
1927

2028
### Changed
2129
- `explain_clean()` now profiles only post-clean columns that can contribute a
@@ -34,6 +42,99 @@ adheres to [Semantic Versioning](https://semver.org/).
3442
zero-column frame that keeps the row count, as the pandas pipeline does.
3543
Polars and Arrow output cannot hold rows without columns; the report records
3644
that difference (#201).
45+
- `freshdata validate` exits 2 ("could not load rules") instead of 1 when a
46+
`--suite` or `--contract` file is not an object, and
47+
`ValidationSuite.from_dict` raises `ValueError` for non-mappings (#289).
48+
- `dbt-gate` no longer passes when nothing was gated. A file without a `nodes`
49+
mapping (such as `run_results.json`) raises `ValueError`, `all_passed` is
50+
false when no models were processed, and `--fail` exits 1. Ephemeral and
51+
disabled models are skipped and listed under a new `skipped` summary key
52+
instead of counting as failures (#296, #249).
53+
- Semantic repairs that could rewrite a valid value are now suggested for
54+
review instead of auto-applied: fuzzy cleaning-memory matches, percent values
55+
in `rate`/`ratio` columns whose scale is fractional or unknown, and shape
56+
alignments of unseparated values. Shape alignments whose groups do not match
57+
the template are no longer proposed (#252, #253, #254).
58+
- The EIDR check character (MD-C002) now uses the hybrid ISO 7064 MOD 37,36
59+
system from the EIDR ID Format spec, so published EIDR IDs validate. IDs
60+
whose check character came from the old MOD 37-2 code are flagged, and a `*`
61+
check character is always rejected (#259).
62+
- The finance FIN-003 guard leaves ambiguous DD/MM vs MM/DD dates unresolved
63+
even when a time follows the date, instead of reading them month-first
64+
(#260).
65+
- HIPAA Safe Harbor reports now rest on column evidence. `fd.clean` and
66+
`fd.apply_plan` record the input columns, so identifier columns that
67+
cleaning did not touch are detected without `dataframe=` and reports that
68+
used to pass can fail. Reports with no full column list (synthetic reports,
69+
native backends, streaming and multi-file runs) set
70+
`coverage_verifiable: False`, add a warning and do not pass (#245).
71+
- HIPAA identifier hints match whole column-name tokens, so `ip` no longer
72+
matches `description` or `ship_address`. Hints of four characters or fewer
73+
no longer match inside run-together lowercase names such as `visitdate`;
74+
separated and camelCase forms still match (#283).
75+
- Blocking rules the pandas entity-resolution backend cannot evaluate (`OR`,
76+
comparison operators, literals, arithmetic, `BETWEEN`, parenthesised
77+
predicates, unquoted names with spaces or hyphens) now raise
78+
`EntityResolutionError` instead of returning zero or wrong pairs. Quote such
79+
names, for example `l."first name" = r."first name"` (#237).
80+
- Entity resolution treats NaT and `pd.NA` as missing, so they no longer count
81+
as agreement and records previously merged on missing datetimes can split
82+
(#238).
83+
- `clean_enterprise` raises `ValueError` when `EnterpriseConfig.anonymization`
84+
is non-empty, instead of silently ignoring a setting it does not apply.
85+
`EnterpriseConfig` still accepts the field, and the CLI prints a one-line
86+
error and exits 1 (#247).
87+
- `TrustScoreWeights` rejects NaN and infinite weights with `ValueError`
88+
instead of producing a `nan` trust score (#277).
89+
- Time-series streaming reads numeric timestamp and event-time columns as
90+
epochs instead of 1970 dates, recording the unit in a
91+
`timeseries_timestamp_parse` action. Unparseable timestamps keep their row
92+
and are reported in `coerced_cells`/`coerced_rows` with a warning, and
93+
mixed-offset batches become a UTC-aware column (#227, #250).
94+
- `TimeSeriesCleanConfig(anomaly_window_size=1)` is rejected when the config
95+
is built (#290).
96+
- `cdc_profile` measures freshness against the current UTC time by default,
97+
and naive `now=` and `watermark=` values are read as UTC (#233, part 4).
98+
- A constant baseline column now produces `drift.ks` findings when current
99+
values move away from the constant. Saved baseline JSON stores statistics at
100+
full precision instead of six decimal places; existing files load unchanged.
101+
Some point-mass shifts score lower than before, because the higher scores
102+
came from the tie handling fixed in #234 (#235, #275).
103+
- `load_review_decisions` reads CSV ids as strings (`"007"` stays `"007"`), and
104+
`apply_review_decisions` raises `ValueError` for a decision it cannot
105+
resolve to a pair instead of dropping it. `feedback_summary` gains an
106+
`n_unmatched` count, and clusters created by an apply get ids numbered past
107+
`n_records` (#239, #267, #268).
108+
- Learned clean values in cleaning-memory JSON are plain JSON types: integers
109+
stay numbers, and Timestamps and Decimals are strings (#256).
110+
- With `apply_plan(allow_drift=True)`, actions whose raw value is no longer in
111+
the column are recorded as skipped (`frame drift`) with count 0, and applied
112+
actions record the observed cell count instead of the plan-time
113+
`n_affected` (#258).
114+
- `Parser.read_text` defaults to `encoding="utf-8-sig"` and strips a leading
115+
UTF-8 BOM from text input (#314).
116+
- `mostly` thresholds are inclusive: a rule with exactly the allowed share of
117+
violating rows (for example 1 of 10 under `mostly=0.9`) warns and passes
118+
instead of failing (#307).
119+
- `RepairPlan.decisions_hash` and `to_json` change for plans whose params hold
120+
sets (members are now sorted) or numpy scalars (now JSON numbers and bools),
121+
so the value no longer depends on `PYTHONHASHSEED` or the numpy version.
122+
Plans without such params keep their hash (#312).
123+
- The context compiler no longer splits allowed values or dedup keys on `/`,
124+
and a bare number is no longer read as a confidence gate (`only if 3
125+
neighbours agree`); such phrases stay unparsed and raise under `strict`
126+
(#301, #303).
127+
- `fd.clean(..., policy=..., strict=True)` with a schema-free policy now
128+
raises `protection_conflict`, as the `columns=` and `context=` flows do
129+
(#304).
130+
- Phone validation rejects numbers with more than 15 digits (E.164) (#318).
131+
- Quality-debt `duplicates` now counts duplicate rows detected in the cleaned
132+
output, not only rows removed, so frames with duplicates can warn or fail the
133+
gate under default options. `pii_risk` counts distinct PII columns instead of
134+
matching cells, so scores drop for tall PII columns (#264, #286).
135+
- Network plugins registered without `allow_network=True` re-read
136+
`FRESHDATA_ALLOW_NETWORK_PLUGINS` at call time, so setting the variable after
137+
registration activates them and unsetting it deactivates them again (#299).
37138

38139
### Fixed
39140
- Trust-gate integrations now validate `on_low_score` policies at configuration
@@ -98,6 +199,104 @@ adheres to [Semantic Versioning](https://semver.org/).
98199
- The missing-pyarrow error names the feature that needs it (for example Arrow
99200
output), and Parquet metadata reads no longer fail with `AttributeError` in a
100201
fresh process (#215).
202+
- `freshdata clean --config` reports invalid YAML or JSON, non-object
203+
sections and unknown keys as a one-line error naming the file (exit 1)
204+
instead of a traceback, and `dbt-gate` does the same for malformed manifests
205+
and directory paths (#289).
206+
- `freshdata clean` and `freshdata validate` no longer exit 1 after a passed
207+
gate when stdout cannot encode UTF-8 (#295).
208+
- Semantic memory replay looks up the expert that learned a repair, so learned
209+
Unicode normalization and shape-alignment repairs replay instead of being
210+
checked as dates (#300).
211+
- Retail GTIN checks read float-loaded integral cells as integer text, so a
212+
blank cell no longer causes a valid GTIN to be rewritten into a different
213+
one (#229).
214+
- Healthcare date checks compare offset-aware FHIR `dateTime` values with naive
215+
dates or mixed offsets in UTC instead of raising `TypeError` (#233, part 2).
216+
- The GDPR Article 30 report lists only the measures the run actually applied
217+
(#287).
218+
- The DuckDB entity-resolution backend accepts non-equality blocking SQL such
219+
as `jaro_winkler_similarity(l.name, r.name) > 0.8` (#236).
220+
- `fd.link(backend="duckdb")` works on keys containing spaces or hyphens
221+
(#266).
222+
- `link_entities` and external `fd.link` reports record their thresholds, so
223+
`build_review_queue` orders items around the configured midpoint (#271).
224+
- `StreamingCleaner(global_duplicates=True)` keeps a bounded window of recent
225+
rows instead of the first `window_size` rows forever, so duplicates of recent
226+
rows are removed for the whole stream (#292).
227+
- Streaming cross-batch deduplication no longer misses duplicates when a
228+
column flips between integer and float dtypes (#293).
229+
- Streaming distribution drift fires for a column that had been constant and
230+
then changes (#294).
231+
- MAD anomaly detection no longer flags the minority value of two-valued or
232+
sparse series (#291), and anomaly columns stay stable when a batch's dtype
233+
changes, with text cells never scored or capped (#248, part).
234+
- CSV review queues round-trip: the formula guard added on export is removed
235+
from ids on load, blank decision cells are skipped, and applying decisions
236+
keeps existing cluster ids and canonical records (#239, #240, #268).
237+
- A frame no longer fails drift checks against its own baseline when values
238+
tie across stored quantiles (#234).
239+
- `pd.ArrowDtype` columns (double, decimal, large string, dictionary) are
240+
recognised by contracts and baselines, and Arrow decimals no longer crash
241+
numeric profiling (#241).
242+
- Semantic cross-field checks no longer raise or miss findings on frames with
243+
duplicate row labels (#231, part 1), integer column labels no longer raise
244+
`KeyError` in semantic repair (#232, part 1), and date-ordering checks
245+
compare tz-aware and naive values instead of raising `TypeError` (#233,
246+
part 1).
247+
- `save_profile` no longer fails with `TypeError` on profiles learned from
248+
numpy or pandas clean values (#256), and `LearningProfile.merge` no longer
249+
modifies either parent's memory (#257).
250+
- The test suite runs from an unpacked sdist and in isolation, and the
251+
streaming docs describe `rolling_trust_score` as an unweighted mean
252+
(#347, #348, #349).
253+
- The FHIR parser records malformed resources (unexpected list or object
254+
shapes, non-string `resourceType`) as per-resource warnings instead of
255+
raising (#313).
256+
- HL7v2, FHIR and EDIFACT parsers handle input that starts with a UTF-8 BOM
257+
(#314).
258+
- The GPX parser skips points with NaN, infinite or out-of-range coordinates
259+
with a warning (#320).
260+
- The context compiler keeps quoted values and values such as
261+
`Trinidad and Tobago` whole (#301), and dotted column names such as
262+
`file.name` no longer split the sentence (#302).
263+
- `clean_text_value` is idempotent and returns text in the configured Unicode
264+
normal form (#316).
265+
- `lint_text_encoding` no longer reports Japanese text such as `コーヒー`, or
266+
`Nº 5`, as mixed script, or uppercase Portuguese such as `MANHÃ DE SOL` as
267+
mojibake (#317).
268+
- `validate_fields` accepts international phone numbers such as
269+
`+49 (0) 30 12345678` and punycode email TLDs (#318).
270+
- `insight_report` issue ids are unique for columns whose names slug the same
271+
(#329).
272+
- HTML report filter boxes work; the generated script was a JavaScript syntax
273+
error (#336).
274+
- `stakeholder_summary` and the per-column view no longer describe preserved
275+
missing values as changes (#337), and no longer claim 100% completeness when
276+
every column was dropped (#339).
277+
- `export_dbt_tests` quotes YAML scalars that would change type or lose
278+
characters on load (dates, `0x1F`, trailing newlines), and floats such as
279+
NaN and infinity load back as floats (#342).
280+
- `OnnxEncoder` no longer fails when several threads trigger the lazy model
281+
load at once (#340).
282+
- `fd.models.pull` times out stalled downloads and rejects a truncated file
283+
instead of installing it (#341).
284+
- The FreshCore backend falls back to pandas for configurations its native
285+
kernels got wrong: `impute="missforest"`, per-column `impute_strategy`,
286+
outlier detection on float columns holding infinity, and mode imputation of
287+
nullable boolean columns. Under `fallback_policy="error"` these runs raise
288+
`FallbackError` (#322, #334, #335).
289+
- The FreshCore adapter reports native duplicate detections and applies
290+
`duplicate_ratio_action` to native drop counts, as the pandas pipeline does
291+
(#323, part).
292+
- The Spark engine renames columns without collisions, honours
293+
`duplicate_keep` and input order when deduplicating, reads float `NaN` as
294+
null with outlier fences from finite values only, and no longer treats
295+
interval columns as numeric (#330, #331, #332, #333).
296+
- A malformed plugin proposal or entry point is dropped or skipped with a log
297+
message instead of crashing the clean or stopping registration (#297).
298+
- Reusing a plugin name within one kind logs a warning naming the replaced
299+
plugin (#298).
101300

102301
## [2.0.0] - 2026-07-20
103302

0 commit comments

Comments
 (0)