@@ -16,6 +16,14 @@ adheres to [Semantic Versioning](https://semver.org/).
1616 tool orderings while keeping PyJanitor optional.
1717- A dependency-optional Great Expectations recipe demonstrating the
1818 repair-then-validate workflow with an in-memory checkpoint.
19+ - ` TimeSeriesCleanConfig.timestamp_unit ` (` "s" ` , ` "ms" ` , ` "us" ` or ` "ns" ` ) sets
20+ the epoch unit for numeric timestamp and event-time columns in time-series
21+ streaming. Without it the unit is inferred from the values (#227 ).
22+ - ` apply_review_decisions() ` accepts a keyword-only ` queue= ` (the
23+ ` ReviewQueueReport ` the reviewer worked from), so decisions that carry only
24+ an ` item_id ` can be resolved to a record pair (#267 ).
25+ - The ` FRESHDATA_MODEL_TIMEOUT ` environment variable sets the network timeout,
26+ in seconds, for ` fd.models.pull ` downloads (default 60) (#341 ).
1927
2028### Changed
2129- ` explain_clean() ` now profiles only post-clean columns that can contribute a
@@ -34,6 +42,99 @@ adheres to [Semantic Versioning](https://semver.org/).
3442 zero-column frame that keeps the row count, as the pandas pipeline does.
3543 Polars and Arrow output cannot hold rows without columns; the report records
3644 that difference (#201 ).
45+ - ` freshdata validate ` exits 2 ("could not load rules") instead of 1 when a
46+ ` --suite ` or ` --contract ` file is not an object, and
47+ ` ValidationSuite.from_dict ` raises ` ValueError ` for non-mappings (#289 ).
48+ - ` dbt-gate ` no longer passes when nothing was gated. A file without a ` nodes `
49+ mapping (such as ` run_results.json ` ) raises ` ValueError ` , ` all_passed ` is
50+ false when no models were processed, and ` --fail ` exits 1. Ephemeral and
51+ disabled models are skipped and listed under a new ` skipped ` summary key
52+ instead of counting as failures (#296 , #249 ).
53+ - Semantic repairs that could rewrite a valid value are now suggested for
54+ review instead of auto-applied: fuzzy cleaning-memory matches, percent values
55+ in ` rate ` /` ratio ` columns whose scale is fractional or unknown, and shape
56+ alignments of unseparated values. Shape alignments whose groups do not match
57+ the template are no longer proposed (#252 , #253 , #254 ).
58+ - The EIDR check character (MD-C002) now uses the hybrid ISO 7064 MOD 37,36
59+ system from the EIDR ID Format spec, so published EIDR IDs validate. IDs
60+ whose check character came from the old MOD 37-2 code are flagged, and a ` * `
61+ check character is always rejected (#259 ).
62+ - The finance FIN-003 guard leaves ambiguous DD/MM vs MM/DD dates unresolved
63+ even when a time follows the date, instead of reading them month-first
64+ (#260 ).
65+ - HIPAA Safe Harbor reports now rest on column evidence. ` fd.clean ` and
66+ ` fd.apply_plan ` record the input columns, so identifier columns that
67+ cleaning did not touch are detected without ` dataframe= ` and reports that
68+ used to pass can fail. Reports with no full column list (synthetic reports,
69+ native backends, streaming and multi-file runs) set
70+ ` coverage_verifiable: False ` , add a warning and do not pass (#245 ).
71+ - HIPAA identifier hints match whole column-name tokens, so ` ip ` no longer
72+ matches ` description ` or ` ship_address ` . Hints of four characters or fewer
73+ no longer match inside run-together lowercase names such as ` visitdate ` ;
74+ separated and camelCase forms still match (#283 ).
75+ - Blocking rules the pandas entity-resolution backend cannot evaluate (` OR ` ,
76+ comparison operators, literals, arithmetic, ` BETWEEN ` , parenthesised
77+ predicates, unquoted names with spaces or hyphens) now raise
78+ ` EntityResolutionError ` instead of returning zero or wrong pairs. Quote such
79+ names, for example ` l."first name" = r."first name" ` (#237 ).
80+ - Entity resolution treats NaT and ` pd.NA ` as missing, so they no longer count
81+ as agreement and records previously merged on missing datetimes can split
82+ (#238 ).
83+ - ` clean_enterprise ` raises ` ValueError ` when ` EnterpriseConfig.anonymization `
84+ is non-empty, instead of silently ignoring a setting it does not apply.
85+ ` EnterpriseConfig ` still accepts the field, and the CLI prints a one-line
86+ error and exits 1 (#247 ).
87+ - ` TrustScoreWeights ` rejects NaN and infinite weights with ` ValueError `
88+ instead of producing a ` nan ` trust score (#277 ).
89+ - Time-series streaming reads numeric timestamp and event-time columns as
90+ epochs instead of 1970 dates, recording the unit in a
91+ ` timeseries_timestamp_parse ` action. Unparseable timestamps keep their row
92+ and are reported in ` coerced_cells ` /` coerced_rows ` with a warning, and
93+ mixed-offset batches become a UTC-aware column (#227 , #250 ).
94+ - ` TimeSeriesCleanConfig(anomaly_window_size=1) ` is rejected when the config
95+ is built (#290 ).
96+ - ` cdc_profile ` measures freshness against the current UTC time by default,
97+ and naive ` now= ` and ` watermark= ` values are read as UTC (#233 , part 4).
98+ - A constant baseline column now produces ` drift.ks ` findings when current
99+ values move away from the constant. Saved baseline JSON stores statistics at
100+ full precision instead of six decimal places; existing files load unchanged.
101+ Some point-mass shifts score lower than before, because the higher scores
102+ came from the tie handling fixed in #234 (#235 , #275 ).
103+ - ` load_review_decisions ` reads CSV ids as strings (` "007" ` stays ` "007" ` ), and
104+ ` apply_review_decisions ` raises ` ValueError ` for a decision it cannot
105+ resolve to a pair instead of dropping it. ` feedback_summary ` gains an
106+ ` n_unmatched ` count, and clusters created by an apply get ids numbered past
107+ ` n_records ` (#239 , #267 , #268 ).
108+ - Learned clean values in cleaning-memory JSON are plain JSON types: integers
109+ stay numbers, and Timestamps and Decimals are strings (#256 ).
110+ - With ` apply_plan(allow_drift=True) ` , actions whose raw value is no longer in
111+ the column are recorded as skipped (` frame drift ` ) with count 0, and applied
112+ actions record the observed cell count instead of the plan-time
113+ ` n_affected ` (#258 ).
114+ - ` Parser.read_text ` defaults to ` encoding="utf-8-sig" ` and strips a leading
115+ UTF-8 BOM from text input (#314 ).
116+ - ` mostly ` thresholds are inclusive: a rule with exactly the allowed share of
117+ violating rows (for example 1 of 10 under ` mostly=0.9 ` ) warns and passes
118+ instead of failing (#307 ).
119+ - ` RepairPlan.decisions_hash ` and ` to_json ` change for plans whose params hold
120+ sets (members are now sorted) or numpy scalars (now JSON numbers and bools),
121+ so the value no longer depends on ` PYTHONHASHSEED ` or the numpy version.
122+ Plans without such params keep their hash (#312 ).
123+ - The context compiler no longer splits allowed values or dedup keys on ` / ` ,
124+ and a bare number is no longer read as a confidence gate (`only if 3
125+ neighbours agree` ); such phrases stay unparsed and raise under ` strict`
126+ (#301 , #303 ).
127+ - ` fd.clean(..., policy=..., strict=True) ` with a schema-free policy now
128+ raises ` protection_conflict ` , as the ` columns= ` and ` context= ` flows do
129+ (#304 ).
130+ - Phone validation rejects numbers with more than 15 digits (E.164) (#318 ).
131+ - Quality-debt ` duplicates ` now counts duplicate rows detected in the cleaned
132+ output, not only rows removed, so frames with duplicates can warn or fail the
133+ gate under default options. ` pii_risk ` counts distinct PII columns instead of
134+ matching cells, so scores drop for tall PII columns (#264 , #286 ).
135+ - Network plugins registered without ` allow_network=True ` re-read
136+ ` FRESHDATA_ALLOW_NETWORK_PLUGINS ` at call time, so setting the variable after
137+ registration activates them and unsetting it deactivates them again (#299 ).
37138
38139### Fixed
39140- Trust-gate integrations now validate ` on_low_score ` policies at configuration
@@ -98,6 +199,104 @@ adheres to [Semantic Versioning](https://semver.org/).
98199- The missing-pyarrow error names the feature that needs it (for example Arrow
99200 output), and Parquet metadata reads no longer fail with ` AttributeError ` in a
100201 fresh process (#215 ).
202+ - ` freshdata clean --config ` reports invalid YAML or JSON, non-object
203+ sections and unknown keys as a one-line error naming the file (exit 1)
204+ instead of a traceback, and ` dbt-gate ` does the same for malformed manifests
205+ and directory paths (#289 ).
206+ - ` freshdata clean ` and ` freshdata validate ` no longer exit 1 after a passed
207+ gate when stdout cannot encode UTF-8 (#295 ).
208+ - Semantic memory replay looks up the expert that learned a repair, so learned
209+ Unicode normalization and shape-alignment repairs replay instead of being
210+ checked as dates (#300 ).
211+ - Retail GTIN checks read float-loaded integral cells as integer text, so a
212+ blank cell no longer causes a valid GTIN to be rewritten into a different
213+ one (#229 ).
214+ - Healthcare date checks compare offset-aware FHIR ` dateTime ` values with naive
215+ dates or mixed offsets in UTC instead of raising ` TypeError ` (#233 , part 2).
216+ - The GDPR Article 30 report lists only the measures the run actually applied
217+ (#287 ).
218+ - The DuckDB entity-resolution backend accepts non-equality blocking SQL such
219+ as ` jaro_winkler_similarity(l.name, r.name) > 0.8 ` (#236 ).
220+ - ` fd.link(backend="duckdb") ` works on keys containing spaces or hyphens
221+ (#266 ).
222+ - ` link_entities ` and external ` fd.link ` reports record their thresholds, so
223+ ` build_review_queue ` orders items around the configured midpoint (#271 ).
224+ - ` StreamingCleaner(global_duplicates=True) ` keeps a bounded window of recent
225+ rows instead of the first ` window_size ` rows forever, so duplicates of recent
226+ rows are removed for the whole stream (#292 ).
227+ - Streaming cross-batch deduplication no longer misses duplicates when a
228+ column flips between integer and float dtypes (#293 ).
229+ - Streaming distribution drift fires for a column that had been constant and
230+ then changes (#294 ).
231+ - MAD anomaly detection no longer flags the minority value of two-valued or
232+ sparse series (#291 ), and anomaly columns stay stable when a batch's dtype
233+ changes, with text cells never scored or capped (#248 , part).
234+ - CSV review queues round-trip: the formula guard added on export is removed
235+ from ids on load, blank decision cells are skipped, and applying decisions
236+ keeps existing cluster ids and canonical records (#239 , #240 , #268 ).
237+ - A frame no longer fails drift checks against its own baseline when values
238+ tie across stored quantiles (#234 ).
239+ - ` pd.ArrowDtype ` columns (double, decimal, large string, dictionary) are
240+ recognised by contracts and baselines, and Arrow decimals no longer crash
241+ numeric profiling (#241 ).
242+ - Semantic cross-field checks no longer raise or miss findings on frames with
243+ duplicate row labels (#231 , part 1), integer column labels no longer raise
244+ ` KeyError ` in semantic repair (#232 , part 1), and date-ordering checks
245+ compare tz-aware and naive values instead of raising ` TypeError ` (#233 ,
246+ part 1).
247+ - ` save_profile ` no longer fails with ` TypeError ` on profiles learned from
248+ numpy or pandas clean values (#256 ), and ` LearningProfile.merge ` no longer
249+ modifies either parent's memory (#257 ).
250+ - The test suite runs from an unpacked sdist and in isolation, and the
251+ streaming docs describe ` rolling_trust_score ` as an unweighted mean
252+ (#347 , #348 , #349 ).
253+ - The FHIR parser records malformed resources (unexpected list or object
254+ shapes, non-string ` resourceType ` ) as per-resource warnings instead of
255+ raising (#313 ).
256+ - HL7v2, FHIR and EDIFACT parsers handle input that starts with a UTF-8 BOM
257+ (#314 ).
258+ - The GPX parser skips points with NaN, infinite or out-of-range coordinates
259+ with a warning (#320 ).
260+ - The context compiler keeps quoted values and values such as
261+ ` Trinidad and Tobago ` whole (#301 ), and dotted column names such as
262+ ` file.name ` no longer split the sentence (#302 ).
263+ - ` clean_text_value ` is idempotent and returns text in the configured Unicode
264+ normal form (#316 ).
265+ - ` lint_text_encoding ` no longer reports Japanese text such as ` コーヒー ` , or
266+ ` Nº 5 ` , as mixed script, or uppercase Portuguese such as ` MANHÃ DE SOL ` as
267+ mojibake (#317 ).
268+ - ` validate_fields ` accepts international phone numbers such as
269+ ` +49 (0) 30 12345678 ` and punycode email TLDs (#318 ).
270+ - ` insight_report ` issue ids are unique for columns whose names slug the same
271+ (#329 ).
272+ - HTML report filter boxes work; the generated script was a JavaScript syntax
273+ error (#336 ).
274+ - ` stakeholder_summary ` and the per-column view no longer describe preserved
275+ missing values as changes (#337 ), and no longer claim 100% completeness when
276+ every column was dropped (#339 ).
277+ - ` export_dbt_tests ` quotes YAML scalars that would change type or lose
278+ characters on load (dates, ` 0x1F ` , trailing newlines), and floats such as
279+ NaN and infinity load back as floats (#342 ).
280+ - ` OnnxEncoder ` no longer fails when several threads trigger the lazy model
281+ load at once (#340 ).
282+ - ` fd.models.pull ` times out stalled downloads and rejects a truncated file
283+ instead of installing it (#341 ).
284+ - The FreshCore backend falls back to pandas for configurations its native
285+ kernels got wrong: ` impute="missforest" ` , per-column ` impute_strategy ` ,
286+ outlier detection on float columns holding infinity, and mode imputation of
287+ nullable boolean columns. Under ` fallback_policy="error" ` these runs raise
288+ ` FallbackError ` (#322 , #334 , #335 ).
289+ - The FreshCore adapter reports native duplicate detections and applies
290+ ` duplicate_ratio_action ` to native drop counts, as the pandas pipeline does
291+ (#323 , part).
292+ - The Spark engine renames columns without collisions, honours
293+ ` duplicate_keep ` and input order when deduplicating, reads float ` NaN ` as
294+ null with outlier fences from finite values only, and no longer treats
295+ interval columns as numeric (#330 , #331 , #332 , #333 ).
296+ - A malformed plugin proposal or entry point is dropped or skipped with a log
297+ message instead of crashing the clean or stopping registration (#297 ).
298+ - Reusing a plugin name within one kind logs a warning naming the replaced
299+ plugin (#298 ).
101300
102301## [ 2.0.0] - 2026-07-20
103302
0 commit comments