All notable changes to this project are documented here. The format follows Keep a Changelog, and the project adheres to Semantic Versioning.
-
A repeated index label no longer breaks cleaning.
fd.clean(df)andfd.suggest_plan(df)raised on ordinary input whenever the frame's index carried the same label twice — the shape you get from concatenating two exports withoutreset_index. Two steps selected rows by label, and.locexpands a repeated label to every row carrying it:- the formatted-number rescue in
steps/dtypes.pyassigned withparsed.loc[rescued.index] = rescued.to_numpy(), so the left-hand side was longer than the values and pandas raisedValueError: cannot set using a list-like indexer with a different length than the value; - the coercion-casualty scan in
engine/missing.pynarrowed withdf[col].loc[rows].isna(), where.locreturned more rows than the mask had, raisingIndexError: boolean index did not match indexed array.
Both now work by position, as the undo-log revert and
fieldcheckalready did. The index itself is left alone:steps/duplicates.pyreads it to decide whether repeated timestamps in aDatetimeIndexare real observations rather than duplicate rows, andengine/context.pyreads it too, so replacing it with a positional index would have changed those decisions rather than only the mechanics.Compatibility impact: none for a unique index — that path is unchanged. Frames with repeated labels that previously raised now clean, and produce the same result as the identical frame with a unique index.
- the formatted-number rescue in
-
A time of day no longer defeats the
dayfirst="auto"quarantine. Under the defaultdayfirst="auto", day/month-ambiguous dates are documented to be quarantined for review, never read in an inferred order — and a bare"05/01/2021"was. But"05/01/2021 00:00"was silently read as 1 May and reported as a plainconverted to datetime64[ns], with no quarantine, no ambiguity note and no warning. On a day-first export every date landed in the wrong month. The detector matched only a bareDD/MM/YYYY: its pattern was anchored at$straight after the year, so any time suffix stopped the match and the value fell through to the parser, where"auto"becomes month-first. The detector now searches for the numeric date token anywhere in the value instead of matching the whole string, since no time, weekday, AM/PM marker or UTC offset around it says which part is the month. Searching makes it fail closed: a first fix that listed allowed time suffixes still let"05/01/2021, 10:00","Mon 05/01/2021","05/01/2021.","05/01/2021 10h"and"05/01/2021 at 10:00"through, because any such list is incomplete. Affects every supported pandas version. The finance domain validator already treated timed values as ambiguous; the core dtype step now agrees with it.Compatibility impact: output changes for columns of numeric dates with anything around them (a time, weekday, offset, trailing punctuation) where both day and month are <= 12, cleaned with the default
dayfirst="auto". Such values are now quarantined — set to missing, with the original kept inreport.coerced_cells— exactly as the same dates without a time always were, and a column made entirely of them stays text. Previously they were converted, month-first. To convert them, state the order:dayfirst=Trueordayfirst=False. Unambiguous values ("13/01/2021 10:00") and explicitdayfirstsettings are unchanged. -
dayfirst=Trueno longer reinterprets unambiguous ISO-8601 dates.fd.clean(df, dayfirst=True)on["2021-01-05", "2021-02-11"]silently returned2021-05-01,2021-11-02— month and day swapped, with no warning and no coercion record.dayfirstresolves ambiguous short dates such as05/12/2021; an ISO-8601 date isYYYY-MM-DDby definition and has no ambiguity to resolve. Cause: pandas 2 infers one format for a whole column from its first value and, underdayfirst=True, reads an ISO date as%Y-%d-%m;_parse_datetimenow passesformat="mixed"in that case so each value is read by its own shape. pandas 1.x infers per value and was never affected, and has noformat="mixed", so the change is guarded on the pandas major version.The corruption was data-dependent, which is what made it easy to miss: it was silent only while every day was
<= 12, because a day>= 13made the guessed format fail, dropped the parse share belowdatetime_threshold, and triggered the mixed-format retry that produced the correct reading. The same column therefore read correctly or incorrectly depending on values it happened to contain, or on an unrelated threshold.Compatibility impact: output changes for ISO-8601 columns cleaned with
dayfirst=Trueon pandas 2 — from a wrong reading to the correct one.dayfirstbehaviour on genuinely ambiguous slash dates is unchanged, as are thedayfirst="auto"andsemantic_contextroutes, which were never affected. -
A complex number in an object column no longer crashes cleaning.
fd.clean(pd.DataFrame({"v": [complex(1, 2), "abc", "3"]}))raisedTypeError: ufunc 'remainder' not supported for the input typesfrom the numeric finalizer's integrality check (nonnull % 1 == 0), and so did a complex value among numeric strings or among Python numbers in an object column (Pythoncomplex,numpy.complex128andnumpy.complex64alike). Cause: once one cell is complex,pd.to_numericreturns a complex result, and it never writes a text, bytes or bool cell into that result — those cells come back as whatever numpy's last freed buffer of the same size held."abc"and"3"could read as0j,7+7jor NaN, so the column passed or failednumeric_thresholddepending on what had run earlier in the process, and when it passed the finalizer raised. Every numeric target insteps/dtypes.pyis a real dtype, so a complex value is now treated like any other value the parse cannot interpret: it is masked before parsing, it counts againstnumeric_threshold, and it is never mangled into a real number. When the column still converts, the complex cell is set to missing with its original inreport.coerced_cells, like a word or aFractionin the same position; when it does not, the column is returned untouched and the type-contamination warning names the row.Compatibility impact: only columns holding a complex value change, and every such column previously raised (or, depending on process state, came back unchanged). Columns without a complex value parse exactly as before — the extra step runs only when pandas returns a complex result. A column made entirely of complex values is still declined before parsing and keeps its dtype.
-
The report no longer calls a datetime parse "converted to object". A column of datetime strings carrying different UTC offsets —
"2021-01-05 00:00:00+01:00","2021-01-06 00:00:00+02:00", the ordinary result of a daylight-saving changeover or of merging exports from two regions — is parsed cell by cell into timestamps that each keep their own offset. No singledatetime64dtype can hold several offsets, so pandas returns them in anobjectcolumn. That parse is faithful and lossless, butfix_dtypesbuilt its action description from the resulting dtype alone and recorded it asconverted to object, risk low, which reads as though nothing was converted. A column mixing values with and without an offset took the same path and got the same record. The description now says what happened:parsed to timestamps; kept as object because the values carry different UTC offsets, so no single datetime64 dtype can hold them(or... because the values mix timezone-aware and timezone-naive timestamps, ...), followed by the existing(N unparseable value(s) set to missing)suffix when cells were coerced. The coercion warning for such a column likewise says the values "could not be parsed as timestamps" instead of "as object". Affects every supported pandas version (pandas 1.x deliversdatetime.datetimecells, pandas 2Timestampcells; the record is the same).Compatibility impact: report text only. Cell values, dtypes, risk levels and confidence are unchanged, and so is every other column's
converted to <dtype>description, including single-offset (datetime64[ns, UTC+01:00]) and naive (datetime64[ns]) datetime columns. Code that matched the literal stringconverted to objectfor such a column needs to match the new description.fd.profilestill previews such a column aswould convert to object.
StreamingCleanConfig.window_sizewas documented as sizing "rolling statistics and the rolling trust score". It does neither: its only effect is to bound the cross-batch duplicate window, and the rolling trust score is sized byrolling_trust_window. Docstring corrected; no behaviour change.docs/repair-plans.mdnow states that theFrameSignaturecontent sample is the first 512 rows, so a change beyond row 512 that preserves row count, column names and dtypes is not detected by drift refusal.- The README's "Native Polars DataFrames" section now states what
docs/fallback-matrix.mdalready did: with default options every native engine delegates the whole pipeline to pandas, and the fully native path isstrategy="conservative"withfix_dtypes=False.
- A domain regex rule no longer fails a valid code because its column is
float64. A numeric code column with one blank cell loads from CSV asfloat64, sostr()renders10000266as"10000266.0"and the trailing0reads as an extra digit — every row of a perfectly valid column then failed a digit-only pattern and the domain trust score dropped with it (GS1-008[0-9]{8}and FIN-008[A-Za-z0-9]{4,12}were both affected).retail/validator.pyalready carried_integral_float_textfor exactly this case on the GTIN checks; it now lives indomains.baseasintegral_float_textand the shared rule engine applies it to every regex rule. Genuine decimals, non-finite floats, text and integers are returned unchanged, so a real violation is still reported. fd.clean_excelnow preserves zero padding, asfd.clean_csvalready did.pandas.read_excelinfers types exactly asread_csvdoes, so a cell that the workbook stored as the text"02134"arrived as the integer2134and the padding was gone before any cleaning step ran.clean_csvavoids this with a bounded pre-scan;clean_excelhad no equivalent, sopreserve_leading_zeros=True— documented as a shared option — changed nothing there, and a postcode or account column was silently read as a quantity and then profiled and outlier-checked as one. Aread_excelcounterpart of the pre-scan now reads only the zero-padded numeric columns as text.preserve_leading_zeros=Falsestill opts out, an explicitread_excel_kwargs={"dtype": ...}still wins,sheet_nameis honoured, and columns without padding keep their numeric dtype. Default-output change: a zero-padded numeric column in a spreadsheet now cleans as text rather than losing its padding.- An unrecognised
semantic_typeno longer receives more text cleaning than a recognised one.textclean.config_for_fieldmatched the declared type exactly — case-sensitively and untrimmed — and fell through to the caller's config on no match, so"Ticker"was cleaned more aggressively than"ticker", and a column declared"password"or"api_key"was case-folded and stripped of punctuation while"identifier"was protected. The lookup is now normalised (casefolded and trimmed), and an unrecognised type warns when a lossy option is active, asfieldcheckalready does for an unknownsemantic_type. Only opt-in options (case,remove_punctuation,strip_html,strip_urls,max_char_repeat,max_length) were ever affected, so a defaultfd.cleanis unchanged and the default path stays warning-free. - Currency parsing no longer assumes a US locale for every currency. Every
comma was deleted and the dot was taken as the decimal point regardless of
the currency present, so
"EUR 1.200,50"read as 1.2005 — a thousand-fold error on a monetary amount —"€0,50"read as 50.0, turning fifty cents into fifty euros, and"€1.234.567,89"failed to parse at all. Undersemantic_mode="auto"the repair was applied automatically at confidence 0.98 and risklow. The locale is now resolved deterministically from the value's own punctuation: with both separators present the right-most is the decimal one (so a euro amount written the US way is still read correctly), a repeated separator can only be grouping, and a single separator with a tail that is not three digits must be decimal. Only a single separator followed by exactly three digits is genuinely ambiguous ("1.200"is 1200 in Berlin and 1.2 in Boston), and that is settled by the currency's convention rather than guessed. Malformed grouping such as"$1.2.3"and"$1,20.50"is now rejected instead of being coerced to123and120.5. US and Indian lakh formats are unchanged. Default-output change: a European-format currency string in a money column now cleans to its correct magnitude. - A value the caller has explicitly declared permitted is no longer discarded
as a null marker.
fd.clean'snormalize_sentinelsstep applied the built-in sentinel set unconditionally, so"NA"in an ISO-3166 country column (Namibia) and"None"in a brand column became missing even whenallowed_valuesfor that column listed them.fd.validate_fieldshas honoured the opposite rule since theTestAllowedValuesBeatNullMarkersregression — "when the schema literally allows a value, it is a value, not a missing marker" — so the same declaration was respected by one public API and ignored by another. The only previous escapes were protecting the column outright, which disables every other repair, ornormalize_sentinels=False, which is global. A column's declaredallowed_values(whether passed throughsemantic_contextor compiled from acontext=policy) now removes those tokens from that column's sentinel set, matched casefolded and trimmed. The exemption is scoped to the declaring column, and a column with no declaration is unchanged —"NA"with no vocabulary is still read as missing. engine="duckdb"no longer silently changes temporal values on the fully native path (strategy="conservative",fix_dtypes=False). A nanosecondtimedelta64[ns]column was truncated to DuckDB's microsecondINTERVAL(5nscame back as0) and a timezone-aware datetime column came back in the machine's session time zone at microsecond resolution — both with an emptyfallback_events/backend_differences. Such columns now take the disclosed pandas fallback instead, so the values survive unchanged andfallback_policy="error"can refuse the run.periodandintervalcolumns no longer break native ingestion. DuckDB raisedNotImplementedException: Data type 'period[M]' not recognizedand Polars returned a period as its raw int64 ordinal (2020-01→600) and an interval as a{left, right}struct. Both engines now fall back to the pandas reference for those dtypes, with the reason recorded.- A Polars frame whose integer column holds nulls no longer loses large
integers.
pl.DataFrame.to_pandas()renders such a column asfloat64, which rounds every value a float64 cannot represent, sofd.clean(pl_df)returned9007199254740992for an input of2**53 + 1— on the default engine, because every public entry point reads a Polars source through this conversion. Those columns are now rebuilt from the raw integers plus a null mask, giving the pandas nullable dtype of the same width (Int64,UInt64,Int32, …). Integer columns without nulls are untouched. Default-output change: a Polars input whose integer column has nulls now cleans as a nullable integer column instead offloat64, exactly as the same data already did when passed in as pandas. CleanReport.revert()no longer writes a restored value into other rows that share a duplicate index label. The undo log now records positional offsets and revert restores by position, so a frame with a non-unique index reverts exactly to its input instead of overwriting rows that never held the value (older reports without positions still revert best-effort by label).- The native Polars backend no longer raises on polars versions that reject
collect(engine="streaming")with aValueError(polars 1.1–1.24): the streaming collect now falls back to a plaincollect()on those versions, as it already did for the olderstreaming=keyword. Modern polars is unaffected. - Explicit imputation (
impute="mean","median","mode","auto","missforest"or animpute_strategyentry) no longer fills the declaredid_columnsortarget_column, as documented. Those columns keep their missing values and the report recordsskipped: identifier columnorskipped: target column. Animpute_strategyentry naming one of them is ignored with a warning. Declared names resolve after column renaming, and the columns can still serve as MissForest features for other columns. fd.cleanon a Spark DataFrame no longer raisesTypeError: cannot materialize source of type DataFrame. Under the defaultstrategy="balanced"the pandas fallback now materializes a Spark source through itstoPandas()method, alongside the existing polars and DuckDB paths.fd.evaluate_quality_debtno longer scores a dimension a clean 0.0 when it was never measured.type_instabilitywhen profiling fails,pii_riskwhen the PII scan is unavailable or fails, andschema_driftandcategory_churnwhen nobaseline=is passed are now not assessed, using the same mechanism as undetectable duplicates (#414):scoreandover_thresholdserialise asNone, the detail says why, the item is listed inQualityDebtGate.unassessed, and it counts toward neither the total, the gate status nor the ledger history. Runs with a baseline and a working profile and PII scan score exactly as before.compute_trust_scoreno longer flags an object column as "mixed types", or lowers consistency for it, when every non-null value is a list (or every one a dict, or every one a tuple). Such a column now scores like the equivalent nested Arrow column. Columns that mix kinds, such as strings with numbers or lists with scalars, are still flagged.StreamingCleanerno longer raisesTypeError: Cannot compare tz-naive and tz-aware timestampswhen a datetime column is tz-naive in one batch and tz-aware in another (including a naive batch followed by strings with mixed UTC offsets). Running datetime statistics and time-series watermarks compare in UTC, reading naive values as UTC, and keep the zone each value arrived in. Time-series mode warns when a timestamp or event-time column changes awareness between batches, andstate_marks such columnstz_mixed.cdc_profileno longer reads small numbers such as row numbers1..100as epoch seconds and reports a 56-year-old batch as passing. When the epoch unit is inferred and the median event time falls before 1990-01-01, it reports anevent_time_implausibleerror and leavesfreshness_secondsunset, without evaluating lateness or ordering. An explicitevent_time_unit=is trusted as before, and epoch columns in s, ms, us and ns from 1990 on are unchanged. Numeric epoch event times before 1990-01-01 now need an explicitevent_time_unit; without it such a batch fails withevent_time_implausible. Time-series streaming warns about such columns whentimestamp_unitis not set.
- GPX and SDMX parsing detects the document encoding (BOM, UTF-16/UTF-32 prefixes, XML declaration) and checks with expat before reading, so DTD and entity declarations are rejected in every encoding. Previously a UTF-16 document bypassed the check and allowed entity expansion (GHSA-2pcm-99rf-3fqq).
- CSV formula sanitising now covers every level of multi-row headers, and index labels and names, so crafted header cells from the input are no longer written as live formulas (GHSA-h3vg-9xq5-7xh4).
- The DuckDB engine no longer spills to the shared
/tmp/freshdata_spill.EngineConfig.temp_directorydefaults toNone, and each run spills into a private (0700), per-run directory under the user's cache directory (orFRESHDATA_SPILL_DIR), which is removed afterwards. An explicittemp_directoryis checked for ownership and permissions (GHSA-q4wj-xrq5-gvww). - Baseline category labels are no longer unkeyed SHA-1. Without
label_keya baseline stores a label-free frequency profile; withlabel_key(orFRESHDATA_BASELINE_KEY) labels are HMAC-SHA256. Baselines are written as schemafreshdata-baseline-v2; v1 baselines still load with a warning and should be rebuilt (GHSA-2826-rpcg-97gg). JsonTokenVaultandSqliteTokenVaultcreate their files owner-only (0600) at creation time. SQLite journal/WAL files inherit that mode. Existing group/other-readable vault files trigger a warning (GHSA-jq8x-9w3v-7j4g).tokenize,surrogateand keylessfpemasking rules, and the policypseudonymizeaction (the default in the GDPR, HIPAA and FERPA packs), no longer fall back to public constants when no key is set. They use a random per-call key and emitEphemeralKeyWarning; passkey=/key_env=for stable, joinable output (GHSA-w58x-9xfc-pq9p).fd.learn(privacy='mask')treats every PII typedetect_piireports (payment cards, IBANs, IP addresses, health and licence identifiers) as sensitive; unknown types fail closed. It adds card and bank column-name hints.freshdata profile auditflags raw card numbers and IBANs in existing profiles (GHSA-hhwr-8mxc-5j9r).clean_enterprisereports no longer contain raw values of masked columns: cluster canonical, variant and key values, semantic-validation invalid samples, andclean_report.coerced_cellsoriginals (and the coercion warnings quoting them) for masked columns are masked or redacted. The[SENSITIVE:xxxxxxxx]tokens that stand in for declaredsensitive_columnsvalues are now a truncated HMAC-SHA256 under a random per-process key instead of an unkeyed SHA-256; they still match within a run but differ between runs (GHSA-hfxp-fcx6-gj93).- Copilot sample masking uses an allow-list: only numeric and boolean sample values pass through, so Arrow-backed string, dictionary and other non-numeric columns (including datetimes) are hash-masked (GHSA-vfqx-wprp-wwg2).
- Copilot masks sample values by column position, so integer, float and tuple
column labels no longer bypass masking.
sensitive_columnsandmust_maskmatch non-string labels, unknownsensitive_columnsraise, labels that collide as strings raise, and masking fails closed (GHSA-239m-28fp-7fh2).
fd.clean_excel(), the Excel companion tofd.clean_csv(): reads one sheet, cleans it, and optionally writes the result, with formula sanitization on by default. Needs the newexcelextra (openpyxl).CleanReport.to_json()andCleanReport.write_json()for first-class audit report serialization without manualjson.dumps(...)calls.- Added a runnable PyJanitor interoperability example that demonstrates both tool orderings while keeping PyJanitor optional.
- A dependency-optional Great Expectations recipe demonstrating the repair-then-validate workflow with an in-memory checkpoint.
TimeSeriesCleanConfig.timestamp_unit("s","ms","us"or"ns") sets the epoch unit for numeric timestamp and event-time columns in time-series streaming. Without it the unit is inferred from the values (#227).apply_review_decisions()accepts a keyword-onlyqueue=(theReviewQueueReportthe reviewer worked from), so decisions that carry only anitem_idcan be resolved to a record pair (#267).- The
FRESHDATA_MODEL_TIMEOUTenvironment variable sets the network timeout, in seconds, forfd.models.pulldownloads (default 60) (#341). cdc_profileacceptsevent_time_unit("s","ms","us"or"ns") to set the epoch unit for numeric event-time columns. Without it the unit is inferred from the values (#327).FreshDataDbtTransformacceptsaudit_name, the stem of its audit file, andrun()accepts a keyword-onlyraise_on_fail. Withraise_on_fail=Falsea failing gate returns its result instead of raising (#343, #344).load_cleaning_memoryaccepts a keyword-onlydataset_idthat selects one memory from a SQLite store holding several (#306).ModelConfig.file_sha256pins a checksum for each file of a model, alongside the existingsha256pin for the primary file (#346).
explain_clean()now profiles only post-clean columns that can contribute a decision narrative, avoiding a redundant full-width context pass.- The
polars,outofcore,enterprise,allanddevextras now requirepolars>=1.0. The Polars engine usesLazyFrame.collect_schema(), which 0.20.x lacks (#214). - The Polars engine reads float
NaNfrom native Polars, Arrow and file sources as null, as pandas input already was. Polars output for those sources shows null where it used to showNaN(#200). - Requesting another engine's native handle (for example
engine="duckdb"withoutput_format="polars-lazy") now raisesValueErrorinstead of returning a different type.engine="auto", or noengine, picks the engine that owns the format (#205). - When every column is dropped as empty, the DuckDB and Polars engines return a zero-column frame that keeps the row count, as the pandas pipeline does. Polars and Arrow output cannot hold rows without columns; the report records that difference (#201).
freshdata validateexits 2 ("could not load rules") instead of 1 when a--suiteor--contractfile is not an object, andValidationSuite.from_dictraisesValueErrorfor non-mappings (#289).dbt-gateno longer passes when nothing was gated. A file without anodesmapping (such asrun_results.json) raisesValueError,all_passedis false when no models were processed, and--failexits 1. Ephemeral and disabled models are skipped and listed under a newskippedsummary key instead of counting as failures (#296, #249).- Semantic repairs that could rewrite a valid value are now suggested for
review instead of auto-applied: fuzzy cleaning-memory matches, percent values
in
rate/ratiocolumns whose scale is fractional or unknown, and shape alignments of unseparated values. Shape alignments whose groups do not match the template are no longer proposed (#252, #253, #254). - The EIDR check character (MD-C002) now uses the hybrid ISO 7064 MOD 37,36
system from the EIDR ID Format spec, so published EIDR IDs validate. IDs
whose check character came from the old MOD 37-2 code are flagged, and a
*check character is always rejected (#259). - The finance FIN-003 guard leaves ambiguous DD/MM vs MM/DD dates unresolved even when a time follows the date, instead of reading them month-first (#260).
- HIPAA Safe Harbor reports now rest on column evidence.
fd.cleanandfd.apply_planrecord the input columns, so identifier columns that cleaning did not touch are detected withoutdataframe=and reports that used to pass can fail. Reports with no full column list (synthetic reports, native backends, streaming and multi-file runs) setcoverage_verifiable: False, add a warning and do not pass (#245). - HIPAA identifier hints match whole column-name tokens, so
ipno longer matchesdescriptionorship_address. Hints of four characters or fewer no longer match inside run-together lowercase names such asvisitdate; separated and camelCase forms still match (#283). - Blocking rules the pandas entity-resolution backend cannot evaluate (
OR, comparison operators, literals, arithmetic,BETWEEN, parenthesised predicates, unquoted names with spaces or hyphens) now raiseEntityResolutionErrorinstead of returning zero or wrong pairs. Quote such names, for examplel."first name" = r."first name"(#237). - Entity resolution treats NaT and
pd.NAas missing, so they no longer count as agreement and records previously merged on missing datetimes can split (#238). clean_enterpriseraisesValueErrorwhenEnterpriseConfig.anonymizationis non-empty, instead of silently ignoring a setting it does not apply.EnterpriseConfigstill accepts the field, and the CLI prints a one-line error and exits 1 (#247).TrustScoreWeightsrejects NaN and infinite weights withValueErrorinstead of producing anantrust score (#277).- Time-series streaming reads numeric timestamp and event-time columns as
epochs instead of 1970 dates, recording the unit in a
timeseries_timestamp_parseaction. Unparseable timestamps keep their row and are reported incoerced_cells/coerced_rowswith a warning, and mixed-offset batches become a UTC-aware column (#227, #250). TimeSeriesCleanConfig(anomaly_window_size=1)is rejected when the config is built (#290).cdc_profilemeasures freshness against the current UTC time by default, and naivenow=andwatermark=values are read as UTC (#233, part 4).- A constant baseline column now produces
drift.ksfindings when current values move away from the constant. Saved baseline JSON stores statistics at full precision instead of six decimal places; existing files load unchanged. Some point-mass shifts score lower than before, because the higher scores came from the tie handling fixed in #234 (#235, #275). load_review_decisionsreads CSV ids as strings ("007"stays"007"), andapply_review_decisionsraisesValueErrorfor a decision it cannot resolve to a pair instead of dropping it.feedback_summarygains ann_unmatchedcount, and clusters created by an apply get ids numbered pastn_records(#239, #267, #268).- Learned clean values in cleaning-memory JSON are plain JSON types: integers stay numbers, and Timestamps and Decimals are strings (#256).
- With
apply_plan(allow_drift=True), actions whose raw value is no longer in the column are recorded as skipped (frame drift) with count 0, and applied actions record the observed cell count instead of the plan-timen_affected(#258). Parser.read_textdefaults toencoding="utf-8-sig"and strips a leading UTF-8 BOM from text input (#314).mostlythresholds are inclusive: a rule with exactly the allowed share of violating rows (for example 1 of 10 undermostly=0.9) warns and passes instead of failing (#307).RepairPlan.decisions_hashandto_jsonchange for plans whose params hold sets (members are now sorted) or numpy scalars (now JSON numbers and bools), so the value no longer depends onPYTHONHASHSEEDor the numpy version. Plans without such params keep their hash (#312).- The context compiler no longer splits allowed values or dedup keys on
/, and a bare number is no longer read as a confidence gate (only if 3 neighbours agree); such phrases stay unparsed and raise understrict(#301, #303). fd.clean(..., policy=..., strict=True)with a schema-free policy now raisesprotection_conflict, as thecolumns=andcontext=flows do (#304).- Phone validation rejects numbers with more than 15 digits (E.164) (#318).
- Quality-debt
duplicatesnow counts duplicate rows detected in the cleaned output, not only rows removed, so frames with duplicates can warn or fail the gate under default options.pii_riskcounts distinct PII columns instead of matching cells, so scores drop for tall PII columns (#264, #286). - Network plugins registered without
allow_network=Truere-readFRESHDATA_ALLOW_NETWORK_PLUGINSat call time, so setting the variable after registration activates them and unsetting it deactivates them again (#299). - Cleaning raises
ValueErrorwhen animpute_strategykey, including one set byPipeline.impute(columns=), names no column after renaming, instead of silently imputing nothing. The message lists the unknown keys and available columns and suggests the normalized name for a pre-rename key such asAge. Keys for columns that a later step drops are still accepted (#310). ExplainReport.cell_changesis always keyed bystr(label).explain_cleanandinfer_rolesraiseValueErrornaming duplicated column labels instead ofAttributeErrororTypeError, andexplain_cleanalso raises for labels whose string forms collide, such as1and"1"(#232, part 3; #265, part 3).suggest_join_keysscores a field 0 when either value is missing (None, NaN, NaT,pd.NAor"") and leaves missing values out of exact-key overlap. Integer keys and NaN-promoted float keys render alike (101and101.0), so they overlap and share blocks (#272, #273).is_valid_icpnrejects values with surrounding text or separators other than spaces and hyphens, such as"tel: 036000291452". Formatted UPC/EAN values such as0-36000-29145-2still pass (#321).cdc_profilereads numeric event-time columns such as Debeziumts_msas epochs in an inferred unit, asclean_timeseriesdoes, instead of as nanoseconds that gave 1970 dates. Small integers such as row numbers are read as seconds (#327).cdc_profilechecks rows with a null CDC key for ordering as one group and counts them in a newmissing_keywarning, which does not affectpassedor penalties. A negativestale_afterraisesValueError(#325, #326).build_baseline,compare_to_baseline,enforce_contractanddiff_schemaraiseValueErrorfor duplicate column labels or labels that collide as strings, and so dofd.validate(suite=...),fd.clean(contract=...)and the enterprise drift step (#265, part 2).- On pandas 2,
min_datetime/max_datetimecontracts parse string columns with mixed date formats instead of dropping values in a second format asNaT. Values that still fail to parse produce a warning-levelcontract.unparseable_datetimefinding, and contracts that passed can now fail (#242). FreshDataDbtTransform(on_low_score="fail").run()raisesTrustGateErrorwhen the gate fails, after writing the audit file, instead of returning a result withshould_fail=True.gate_manifeststill records failing models and gates every model (#343).freshdata clean,freshdata streamandfd.clean_csvload numeric-looking CSV columns with zero-padded values as text, so02134keeps its leading zero. Detection reads a 10,000-row sample (the first chunk when streaming) and is skipped whenread_csv_kwargssetsdtypeorconverters, or withpreserve_leading_zeros=False(#228).freshdata streamexits 1 without writing output when a later batch adds a column or cannot be cast to the first batch's Parquet schema (#248, part a).load_cleaning_memoryon a SQLite store raisesValueErrorlisting the stored ids when the store holds several memories and nodataset_idis given,KeyErrorfor an unknowndataset_id, andFileNotFoundErrorfor a missing path instead of creating an empty database (#306).- Loading a
.fdprofileraisesProfileFormatErrorfor a truncated or malformed manifest or member, and for a manifest that does not cover every required member (#311). - Model checksum pins cover every file. Once a model has a pin, every file
needs one, or
ModelChecksumErrornames the unpinned file before any download.fd.models.pullalso checks files that are already installed and raises on a mismatch with aforce=Truehint (freshdata models pullexits 2), andstatus()reportsverified: Trueonly when every file matches. Models without pins behave as before (#346). - Rewriting
24:00to00:00in time-only columns is suggested for review in every semantic mode instead of auto-applied, and is not learned into cleaning memory. The TruthBench logistics oracle moves log-06transport_time24:00from REPAIR to REVIEW and adds log-07tracking_statuson timetoon-timeas the domain's exact repair (#305). anonymize,detect_piiandapply_privacy_policyleavepd.NAandNaTcells missing instead of rewriting them as"<NA>"or"NaT", andcells_changedno longer counts them (#243).detect_piiraisesValueErrornaming duplicated column labels instead ofAttributeError, and so doesanonymizewhen detection is enabled or a rule targets a duplicated label (#265, part 1).detect_piisetsmetadata["ner"]from whether NER actually ran, and newner_requested,ner_activeandner_errorkeys report the request and outcome. When the Presidio analyzer cannot start, oneUserWarningis emitted, the NER pass is skipped and start-up is not retried within the process (#282).- When
fpecells used more than one mode,metadata["fpe_mode"]is"mixed"andmetadata["fpe_modes"]maps each column to its per-mode cell counts. Single-mode runs report as before (#281, parts 1-2). apply_privacy_policyandclassify_columnsclassify each column from every distinct non-null value instead of the first 200 cells, so columns longer than 200 rows may now be classified. The report metadata gainsclassification_values_scanned(#246).- An in-scope inline
PrivacyPolicyrule now takes priority over every pack rule; classifier specificity still decides within each group (#284). - A policy rule's own key is resolved before the policy defaults:
rule.key_env, thenrule.key, thenpolicy.key_env, thenpolicy.key. Rules without a key resolve as before (#285). apply_privacy_policyandclassify_columnsraiseValueErrorfor duplicate column labels or labels that collide as strings, such as1and"1"(#232, part 5).- A
MaskingRulecolumn name matches every column with the same name or the same snake_case form, so"First Name"masksfirst_nameand"email"masks bothemailandEmail. Listed names that match no column are recorded under a newMaskReport.unmatched_columnskey; they raiseValueErroronly withMaskingRule(strict=True)ormask_dataframe(..., strict=True).freshdata clean --mask COLUMN:STRATEGYrules are strict, so it exits 1 with a one-line error whenCOLUMNmatches no column, instead of masking nothing and exiting 0 (#251).
fd.evaluate_quality_debtreports theduplicatesdimension as not assessed when duplicate rows cannot be checked (object columns of lists or dicts, nested Arrow list/struct/map columns), instead of scoring it a clean 0.0 with "0 duplicate row(s) detected". The item serialises withscoreandover_thresholdasNone, its detail names the unhashable columns, it is listed inQualityDebtGate.unassessed, and it never counts toward the total, the gate status or the ledger history. Frames where duplicates can be detected score exactly as before.compute_trust_scorereports uniqueness as unknown when duplicate rows cannot be checked (the same unhashable columns), instead of a perfect 100.TrustScore.uniquenessisNaNwithuniqueness_assessedFalse,to_dict()givesNone,str()and the Markdown tables shown/a, the blocking columns carry an "unhashable values: duplicate rows not checked" issue, andoverallblends the other three dimensions with their weights renormalised. Frames where duplicates can be detected score exactly as before.- Trust-gate integrations now validate
on_low_scorepolicies at configuration boundaries, rejecting typos instead of silently skipping failure handling (#345). QualityReport.to_markdown()now escapes every cell in the Actions table, so a column name (or description) containing|or a newline no longer adds or splits table columns. Pipes become\|and line breaks become<br>(#338).- The minimum supported numpy is now 1.22. The numpy 1.21.6 wheel bundles an
OpenBLAS that segfaults on BLAS-backed matrix multiplies on current Apple
Silicon Macs regardless of
OPENBLAS_NUM_THREADS, so installs at the old floor could crash in any code path that multiplies float matrices (for example semantic similarity scoring on larger inputs). - Median imputation no longer crashes with
OverflowErroron nullable integer columns (Int8/Int16/Int32/UInt*) containing missing values under numpy 2.5+. Affectsimpute="median"/"auto", the default missing-value engine,fill_missing, MissForest seeding and seasonal time-series imputation. engine="duckdb"withoutput_format="arrow"or"polars"now fetches the result directly in that format instead of materializing a pandas frame withfetchdf()and converting it again (#52). Arrow output keeps DuckDB's column types (for exampleDECIMALstaysdecimal128).- The Polars engine's duplicate-detection and dedup row counts now run through the streaming collect path instead of re-evaluating the whole plan in memory (#53).
fd.clean_timeseriesandStreamingCleanerno longer crash with "cannot convert to 'float64'-dtype NumPy array with missing values" on nullable integer columns (Int*/UInt*) with gaps under pandas < 2 (running statistics and short-gap interpolation). MissForest's convergence check uses the same NA-safe conversion.engine="duckdb"no longer fails with "STDDEV_SAMP is out of range" on float columns holdinginforNaN. Missing counts are exact instead of rebuilt from a rounded percentage, and outlier fences use finite values only (#199).- The Polars engine counts float
NaNas missing, so empty rows and columns are dropped as with pandas, and±infno longer skews outlier fences (#200). - The DuckDB engine really drops all-empty columns when every column is empty, instead of reporting the drop and returning them (#201).
- DuckDB full-row deduplication keeps the input row order and honours
duplicate_keep="first"/"last"(#202). - Polars
LazyFrameinputs work on the default path and when a run falls back to pandas, instead of raisingTypeError(#203). - The DuckDB and Polars engines strip every Unicode whitespace character that
Python's
str.strip()removes (for example NBSP and\x1c–\x1f) (#204). - Pandas inputs with a mixed-type object column or duplicate column labels take a recorded pandas fallback on the DuckDB and Polars engines. They no longer crash (Polars) or silently turn numbers into text (DuckDB) (#206).
CleanReport.revert()no longer writes reverted values into the frame passed to it (#207).- Imputation no longer corrupts nullable integers beyond ±2**53 by casting
them to float64.
impute="mean"/"median", the default engine,fill_missingandStreamingCleanerkeep the integer dtype and present values exact (#208). StreamingCleaner.clean_batchandfd.clean_timeseriesno longer raiseKeyErroron non-string column labels (#209).fd.fill_missingno longer crashes on nullable integer columns with a fractional mean or median, or on nullable boolean columns. Duplicate column labels raise a clearValueError(#210).fd.detect_outliersreturns a plain boolean mask on nullable columns, andfd.remove_outliersno longer drops rows that are only missing.remove_outliers/resolve_duplicateswithinplace=Trueraise on a non-unique index instead of dropping extra rows (#211).pd.ArrowDtypestring columns and categorical text columns get whitespace and sentinel normalization. Categoricals keep their dtype (#212).- Docs: the ydata-profiling comparison example, the CSV sanitization row in trust claims, and the README quickstart output now match current behaviour (#213).
- The missing-pyarrow error names the feature that needs it (for example Arrow
output), and Parquet metadata reads no longer fail with
AttributeErrorin a fresh process (#215). freshdata clean --configreports invalid YAML or JSON, non-object sections and unknown keys as a one-line error naming the file (exit 1) instead of a traceback, anddbt-gatedoes the same for malformed manifests and directory paths (#289).freshdata cleanandfreshdata validateno longer exit 1 after a passed gate when stdout cannot encode UTF-8 (#295).- Semantic memory replay looks up the expert that learned a repair, so learned Unicode normalization and shape-alignment repairs replay instead of being checked as dates (#300).
- Retail GTIN checks read float-loaded integral cells as integer text, so a blank cell no longer causes a valid GTIN to be rewritten into a different one (#229).
- Healthcare date checks compare offset-aware FHIR
dateTimevalues with naive dates or mixed offsets in UTC instead of raisingTypeError(#233, part 2). - The GDPR Article 30 report lists only the measures the run actually applied (#287).
- The DuckDB entity-resolution backend accepts non-equality blocking SQL such
as
jaro_winkler_similarity(l.name, r.name) > 0.8(#236). fd.link(backend="duckdb")works on keys containing spaces or hyphens (#266).link_entitiesand externalfd.linkreports record their thresholds, sobuild_review_queueorders items around the configured midpoint (#271).StreamingCleaner(global_duplicates=True)keeps a bounded window of recent rows instead of the firstwindow_sizerows forever, so duplicates of recent rows are removed for the whole stream (#292).- Streaming cross-batch deduplication no longer misses duplicates when a column flips between integer and float dtypes (#293).
- Streaming distribution drift fires for a column that had been constant and then changes (#294).
- MAD anomaly detection no longer flags the minority value of two-valued or sparse series (#291), and anomaly columns stay stable when a batch's dtype changes, with text cells never scored or capped (#248, part).
- CSV review queues round-trip: the formula guard added on export is removed from ids on load, blank decision cells are skipped, and applying decisions keeps existing cluster ids and canonical records (#239, #240, #268).
- A frame no longer fails drift checks against its own baseline when values tie across stored quantiles (#234).
pd.ArrowDtypecolumns (double, decimal, large string, dictionary) are recognised by contracts and baselines, and Arrow decimals no longer crash numeric profiling (#241).- Semantic cross-field checks no longer raise or miss findings on frames with
duplicate row labels (#231, part 1), integer column labels no longer raise
KeyErrorin semantic repair (#232, part 1), and date-ordering checks compare tz-aware and naive values instead of raisingTypeError(#233, part 1). save_profileno longer fails withTypeErroron profiles learned from numpy or pandas clean values (#256), andLearningProfile.mergeno longer modifies either parent's memory (#257).- The test suite runs from an unpacked sdist and in isolation, and the
streaming docs describe
rolling_trust_scoreas an unweighted mean (#347, #348, #349). - The FHIR parser records malformed resources (unexpected list or object
shapes, non-string
resourceType) as per-resource warnings instead of raising (#313). - HL7v2, FHIR and EDIFACT parsers handle input that starts with a UTF-8 BOM (#314).
- The GPX parser skips points with NaN, infinite or out-of-range coordinates with a warning (#320).
- The context compiler keeps quoted values and values such as
Trinidad and Tobagowhole (#301), and dotted column names such asfile.nameno longer split the sentence (#302). clean_text_valueis idempotent and returns text in the configured Unicode normal form (#316).lint_text_encodingno longer reports Japanese text such asコーヒー, orNº 5, as mixed script, or uppercase Portuguese such asMANHÃ DE SOLas mojibake (#317).validate_fieldsaccepts international phone numbers such as+49 (0) 30 12345678and punycode email TLDs (#318).insight_reportissue ids are unique for columns whose names slug the same (#329).- HTML report filter boxes work; the generated script was a JavaScript syntax error (#336).
stakeholder_summaryand the per-column view no longer describe preserved missing values as changes (#337), and no longer claim 100% completeness when every column was dropped (#339).export_dbt_testsquotes YAML scalars that would change type or lose characters on load (dates,0x1F, trailing newlines), and floats such as NaN and infinity load back as floats (#342).OnnxEncoderno longer fails when several threads trigger the lazy model load at once (#340).fd.models.pulltimes out stalled downloads and rejects a truncated file instead of installing it (#341).- The FreshCore backend falls back to pandas for configurations its native
kernels got wrong:
impute="missforest", per-columnimpute_strategy, outlier detection on float columns holding infinity, and mode imputation of nullable boolean columns. Underfallback_policy="error"these runs raiseFallbackError(#322, #334, #335). - The FreshCore adapter reports native duplicate detections and applies
duplicate_ratio_actionto native drop counts, as the pandas pipeline does (#323, part). - The FreshCore native module counts duplicate rows when
drop_duplicatesis False, at the same stage as the pandas step, soengine="freshcore"records the detection, warns aboveduplicate_thresholdand raisesDuplicateRatioErrorunderduplicate_ratio_action="error"without falling back to pandas (#323). - The Spark engine renames columns without collisions, honours
duplicate_keepand input order when deduplicating, reads floatNaNas null with outlier fences from finite values only, and no longer treats interval columns as numeric (#330, #331, #332, #333). - A malformed plugin proposal or entry point is dropped or skipped with a log message instead of crashing the clean or stopping registration (#297).
- Reusing a plugin name within one kind logs a warning naming the replaced plugin (#298).
- MissForest fills a column that has no predictor columns once, through its
fallback, recording one
missforest_fallbackaction and onecolumns_imputedentry, and no longer needs scikit-learn when every column falls back (#324). fd.validatecomparesallowed_valuesby type on numeric and boolean columns, so1matches1.0andtruematchesTrueinstead of every row being a violation.export_gx_suiteandexport_dbt_testsemit typed numbers and booleans for these sets (#255).explain_cleanreports changed cells, HTML dtypes and narratives for integer-labelled columns, MultiIndex labels no longer breakto_dict()andto_html(), andexplain_cleanandinfer_rolesaccept mixed integer and string labels instead of raisingTypeError(#232, parts 3 and 6).- The HL7 v2 parser uses the delimiters each MSH segment declares, takes the
first repetition for single-valued fields while OBX-5 keeps every repetition
joined with
~, and decodes\F\,\S\,\T\,\R\and\E\escapes (#261). clean_text,validate_fieldsandsuggest_join_keyshandle frames with duplicate row labels, such aspd.concatoutput, instead of writing one row's value to all of them or raising (#231, parts 2-4).validate_fieldscompares tz-aware values with naivemin_valueandmax_valuebounds, or the reverse, in UTC instead of raisingTypeError(#233, part 3).- MissForest keeps integer and nullable integer column dtypes instead of
returning
objectcolumns. Imputed values are rounded half-to-even, noted in the action rationale and flagged by a newrounded_to_integermetadata key (#263). - GTFS-ST004 flags only stop_times rows that repeat an earlier
stop_sequencewithin a trip, so a valid trip whose rows are not sorted no longer raises an error (#319). cdc_profile(stale_after=0)no longer raisesZeroDivisionError, and thereplay_thresholddocs state thatreplay_riskneeds at least one duplicate-key row (#326, #328).diff_schemaaccepts polars frames (#274), and baseline key-level changes treat aNaNkey as one value and accept mixed integer and string keys, so identical frames report no changes (#276).- Contracts and baselines find integer column labels declared as
0or"0"(#232, part 4), and baselines compare tz-aware and naive timestamps in UTC with adrift.timezone_changewarning instead of raisingTypeError(#233, part 5). - The FreshCore backend falls back to pandas for datetime, timedelta,
categorical, period and interval columns, integer columns beyond ±2**53, and
column labels that collide once stringified, instead of changing dtypes or
overwriting columns. Native results restore integer dtypes and non-string
labels; integer columns that cannot be restored stay
float64and are recorded inbackend_differences(#262; #232, part 2). - The privacy extras (
privacyandall) install on Python 3.9 from wheels. On Python 3.9 they cap spaCy below 3.8.8, thinc below 8.3.5 and blis below 1.2.1; Linux aarch64 still builds thinc and blis from source (#278). gate_manifestwithoutput_dirno longer overwrites the audit file of a model that shares its alias with another. Such models write<schema>.<alias>_audit.json, or<unique_id>_audit.jsonwhen the schema is missing or still not unique (#344).freshdata streamoutput keeps the first batch's layout: later CSV batches are written under its columns, so a missing flag column is left empty instead of shifting values, and later Parquet batches are cast to its schema. Output goes to a.partialfile that is moved into place on success, so a failed run leaves no truncated file (#248, part a).- Replaying a decisions table read back from CSV applies table-level steps
such as
drop_duplicates, andCleaningMemory.to_jsonwritesnullinstead of bareNaNorInfinity(#309). - Replaying a merged
LearningProfileon its training frame no longer reports drift or skips text columns, because the merge keeps the parents' source schema (#308). - The GPX and SDMX parsers return their usual invalid-XML warning and empty
frames when the XML declares an unknown encoding, instead of raising
LookupError(#315). check_k_anonymityignores empty combinations of categorical quasi-identifiers, so unused categories no longer givesmallest_class_size=0and fail the check, including throughclean_enterprisewithKAnonymityConfig(#244).- The
reversibleflag infpeaudit metadata follows the mode each cell actually used, so cells that took the surrogate fallback are no longer reported as reversible (#281, parts 1-2). JsonTokenVaultinstances that share a path keep each other's entries: writes re-read and merge the file under a lock instead of rewriting it from a stale copy, and an empty vault file loads as an empty vault.SqliteTokenVaultworks when used from threads other than the one that created it (#279).apply_privacy_policyworks on frames with non-string column labels, such as integers; report keys stay strings (#232, part 5).- The
quarantineprivacy-policy action works on nullable integer, boolean and categorical columns instead of raisingTypeError; those columns come back as object dtype and missing cells stay missing. detect_piiand detection-drivenanonymizescan categorical text columns on pandas 1.5, as on pandas 2 (#280).- Crypto FPE honours
visible(#281). analyze_dataset(mask_salt=...)makesmodel_contextand its fingerprint reproducible; by default they are per-run, andaudit["mask_salt_source"]records which was used (#288).
Remediation of the July 2026 v1.2.0 production-readiness audit: the unsafe defaults it confirmed are now safe-by-default, with every old behavior still available as an explicit opt-in. These default changes are breaking, hence the major version. v1.2.0 was tagged but never published, so PyPI users upgrade directly from 1.1.1 and should read the [1.2.0] section below as part of this release.
drop_duplicatesnow defaults toFalseunder every strategy. Exact duplicate rows are detected and reported but never removed until you opt in withdrop_duplicates=True. A duplicate ratio aboveduplicate_thresholdstill raises the strong report warning; the newduplicate_ratio_action="error"escalates it toDuplicateRatioErrorfor pipelines that must stop on a suspicious join.outlier_action="auto"now flags under every strategy, including"aggressive", and never rewrites values. Winsorizing requires an explicitoutlier_action="cap"; explicitly-requested capping is now skew-aware (log-space Tukey fences for strongly skewed positive data) so legitimate heavy tails are not flattened.dayfirst="auto"never infers a column-wide day/month order. Values whose day-first and month-first readings are both valid are quarantined intoreport.coerced_cells(originals preserved, audit action recorded) unlessdayfirst=True/Falseis set explicitly; a single unambiguous sibling value no longer flips the interpretation of a whole column.- Column-name outlier heuristics are opt-in. The
_DOMAIN_SENSITIVEname match (fraud/amount/risk/…) no longer changes outlier decisions by default; identically-distributed data gets identical treatment regardless of column name. Opt back in withdomain_sensitive_names=True. fd.anonymize()fails closed. Called with norulesand nodetection_configit now raisesValueErrorinstead of returning the data unchanged with a warning.- CSV outputs neutralize spreadsheet formula injection by default
(
fd.clean_csv,freshdata clean/apply-planCLI CSV output, the streaming CLI including its quarantine export, and the HTML-report ledger download). Cells and labels starting with= + - @ <tab> <cr>— including behind leading whitespace — are prefixed with'. Usesanitize_formulas=False/--no-sanitize-formulaswhere byte-exact round-trips matter. - PyPI Development Status classifier downgraded to
4 - Betauntil the project's own absolute release gates (CleanBench T5 runtime gate included) hold on release infrastructure.
- PII detection no longer misreports dates, UUID fragments, licence IDs,
Aadhaar-style 4-4-4 digit groups, semver strings, or ordinary numbers as
PHONE, and no longer misreports IBANs asCREDIT_CARD. Regex matches for PHONE / CREDIT_CARD / IBAN must now pass post-match validators: structural plausibility and boundary checks for phones, Luhn (13–19 digits) for cards, ISO 13616 mod-97 for IBANs (spaced and compact). Checksum-verified findings reportsource="checksum"with score ≥ 0.95. - Restored the nightly perf-regression workflow's CleanBench T5 gate, which
a temporary profiling probe (PRs #152/#153) had left neutered with
|| true— the gate and its failure-alert issue automation are enforcing again. - Refreshed the committed TruthBench
baseline.json, which still recorded the four pre-release-audit KNOWN-RED gates (cleaning:raw_pii_leakage,cleaning:review_routing,cleaning:exact_repair,generated_code_sandbox) even though the release-audit fixes shipped in v1.2.0 made them pass; a regression on any of them now fails the PR ratchet.
Note: v1.2.0 was tagged (
v1.2.0) but never published to PyPI or GitHub Releases — the release was stopped on a benchmark runtime gate. Everything below first shipped to users in 2.0.0.
- TruthBench release runner (
benchmarks/truthbench/): the semantic red-team foundation gained its missing production pieces — an end-to-end runner, normalized decision hashing with fail-closed repeat verification (determinism.py), full generated-code verification (parse → strict AST allowlist → compile → isolatedpython -Iexecution with module poisoning, timeout, input contract, protected-cell and PII-canary checks), a deterministic failure minimizer, atomic schema-validated result artifacts, and a CLI.make truthbench-release(equivalentlyPYTHONPATH=src python -m benchmarks.truthbench run --profile release --backends pandas,polars,duckdb --require-backends --repeats 2 --check) now gates PR CI and the release workflow; every decision/sink surface maps to a concrete behavioral adapter (a contract test rejects placeholders). - Adversarial regression suites for the twelve release-risk hypotheses
(
tests/test_release_hypotheses.py) and the TruthBench components. - Validation Gauntlet (
benchmarks/gauntlet/,docs/validation-gauntlet.md): a gold-labelled disposition benchmark for the validation, domain and text-cleaning surfaces. Five deterministic fixtures (finance, healthcare, CRM, e-commerce, adversarial text) label every injected defect with the disposition FreshData should choose (preserve / repair / flag / review) and the harness scores detection P/R/F1, repair accuracy, review routing, preservation, corruption, escapes, false positives, audit completeness, determinism, trust monotonicity and runtime/memory. Runs on every PR (gauntlet.yml) with absolute gates plus no-regression checks against the storedbaseline.json. CleanReport.coerced_cells: per-cell record ({column: {row: original}}) of values thatfix_dtypesnulled because they did not parse as the column's inferred type — the recovery source for quarantined cells, also included inreport.to_dict().- Date-field range validation in
fd.validate_fields:FieldSpec.min_value/max_valuenow accept a date string or timestamp fordate/datetimefields, so a future date of birth or an 1875 admission date is flagged as adomain_mismatch(gauntlet finding). - Case-variant vocabulary suggestions in
fd.validate_fields: a value that matches anallowed_valuesentry except for case (ACTIVEvsactive) is no longer silently accepted — it gets a warning-severity issue with the canonical form assuggestionand actionaccept_with_warning(gauntlet finding). docs/trust-claims.md(claim-to-evidence map for every trust-relevant README/docs claim, superset of the machine-enforcedCLAIM_REGISTRY) anddocs/threat-model.md(trust boundaries, per-privacy-mode guarantees, ranked residual risks), both linked in the docs nav.benchmarks/bench_outofcore.py: subprocess-isolated peak-RSS evidence for the four engine/output-format combinations on a generated parquet fixture (per-scenarioru_maxrss, wall time, and thematerializedflag).- AI Copilot (experimental) —
freshdata.experimental.ai_copilot.analyze_dataset: deterministic, fully offline dataset analysis that returns a ranked problem list (PII, policy violations, duplicates, missing values, mixed date formats, near-duplicate category spellings), a PII warning, an ordered explainable cleaning plan, and copy-ready freshdata code generated for the analyzed dataset. Privacy-first: raw string values never enter the report'smodel_context(every string-like sample column is hash-masked, numeric values pass through as-is, or samples are omitted entirely withprivacy="schema_only"); the payload is SHA-256 fingerprinted in the audit. An optionalproviderhook (plainCallable[[str], str]) allows plugging in an LLM later — no built-in provider ships, no API key is needed, and provider failures never break the deterministic report. - Flagship demo:
examples/freshdata_ai_copilot_demo.pyplus the bundledexamples/data/messy_customers.csv— the full messy-to-audit-ready story (analyze → mask → clean under a compiled policy → merge category variants → re-score trust), and a new docs guide (docs/ai-copilot.md). CITATION.cffso the project can be cited from GitHub's "Cite this repository" button, and a documentation issue template alongside the existing bug/feature templates.
- Lint now covers the whole repository (
ruff check .in CI, closing #54): benchmark and notebook lint debt fixed, dead code removed (harness_metricsunused gold-labels block), and the ASV-managedfreshdata-benchmarks/sub-project excluded as tool-generated. No runtime behavior changes. - All CI workflows now declare least-privilege
GITHUB_TOKENpermissions (read-only by default; the nightly-alert and coverage-badge jobs keep their scoped write grants). - Contributor docs (
CONTRIBUTING.md,README.md,QUALITY_OPS.md) now quote the exact commands CI runs; the pre-commit config no longer ships aruff-formathook the codebase and CI never enforced.
- Dead packaging/CI config:
MANIFEST.in(ignored by the hatchling build backend — the sdist is shaped bypyproject.toml) andfreshdata-benchmarks/.github/workflows/(nested workflow directories are never executed by GitHub Actions). - Committed AI-assistant working artifacts (
.superpowers/, now git-ignored) and internal planning notes that were being published to the documentation site (docs/superpowers/).
-
Default-path slowdown from the scientific-notation segfault guard (release blocker, nightly issue #147): the guard that masks huge-exponent tokens (
"1e999") beforepd.to_numeric— protection against a pandas 2.3.x segfault — screened text columns cell by cell through a Python predicate, roughly doublingfix_dtypestime on 50k-row frames in CI. Each column is now screened with a single C-level joined-blob regex scan and the per-cell predicate runs only on columns that screen positive. Masking semantics are unchanged; the CleanBench T5 runtime gate is back within its ±20 % baseline envelope. -
dir(freshdata)no longer listsActiontwice: the privacy-policy engine'sActionenum was listed in the lazy enterprise exports but was unreachable there —fd.Actionis (and remains) the audit action fromfreshdata.report. Usefreshdata.enterprise.Actionfor the privacy enum. -
Sensitive-column masking across all report surfaces (
fd.clean(sensitive_columns=...),fd.validate_fields(sensitive_columns=...),analyze_dataset(sensitive_columns=...)): values from declared-sensitive columns never appear verbatim in report warnings, coerced-cell payloads, action rationales/metadata/evidence,normalized_cells, or Copilot-recommended pipelines (which now always mask declared columns before printing report summaries). A deterministic[SENSITIVE:xxxxxxxx]digest token keeps records correlatable without disclosure. -
Ambiguous and partial dates are quarantined, never interpreted: a short-form date whose day/month order cannot be resolved (no explicit
dayfirst, no disambiguating sibling) and partial ISO dates ("2025-01") now coerce to missing with originals preserved incoerced_cellsfor review instead of being silently resolved month-first / to a fabricated day; time-range strings ("09:00-17:00") no longer parse to bogus offset-bearing timestamps. -
Corroboration-gated semantic mutations: unit strips auto-apply only with a declared column unit (inferred units demote to suggestions; unit-mismatched values become high-risk review items), and a new dataset-level
semantic_context["currencies"]declaration routes out-of-policy currency values to review instead of silently dropping the denomination. -
CSV formula injection in
write_exception_table: observed values such as=HYPERLINK(...)were written verbatim to the exception-table CSV and would execute when opened in a spreadsheet. The CSV path now routes through the samesanitize_csv_formulasguard every other spreadsheet export uses. -
Trust-score monotonicity:
compute_trust_scorerated corrupted frames higher than pristine ones because constant columns counted as structural inconsistency (corrupting one made it vary, clearing the flag). Constant columns are now surfaced as per-column issues instead of lowering the consistency dimension. -
Semantic overconfidence: an isolated unit value (
"10 lb"among plain numbers) is no longer auto-stripped at 0.97 confidence — inferred unit consistency now requires majority support and demotes to a suggestion otherwise; already-canonical ISO dates are no longer rewritten to timestamps when unparseable values keep the column as text. -
Calibration (nightly issue #139): the CleanBench full-suite ECE gate failed at 0.0384 > 0.03 from systematic underconfidence of the measured deterministic canonicalization families.
calib-default-2maps their raw 0.96–0.97 scores to measured rule-of-three lower bounds (email_format 148/148, phone_format 168/168, reference_value 128/128 across seeds 0–9) while staying identity below the 0.95 auto threshold, so no apply/suggest/review decision changes. ECE is now 0.0217. -
Default-path performance regression (nightly issues #139/#140): the pandas-segfault exponent guard and the relative-date guard in
fix_dtypesscanned whole columns per value; both are now vectorized (candidate prefilter / unique-first) restoring the T5 runtime gate (slowdown 0.27 → 0.03 vs the v1.0 baseline) with unchanged semantics. -
Nightly online/large lane (issue #138): the lane inherited the repo-wide
--cov-fail-under=93while deliberately selecting ~21 tests, so it could never pass; it now runs with--no-cov(coverage stays enforced on the full fast lane). -
Unparseable values are quarantined, never fabricated (gauntlet finding, the
'apple'-in-a-price-column case): whenfix_dtypesconverts a mostly-numeric (or datetime) text column, cells that fail to parse used to becomeNaNand then be silently imputed by the auto engine — turning junk into a fabricated median. They now stay missing, are excluded from auto-imputation, keep their originals inreport.coerced_cells, and the decision is ahuman_reviewaction in the audit trail. Genuine missing values (trueNaN, sentinels like"N/A") keep the documented auto-impute behaviour, and an explicitimpute=request still fills everything. -
Formatted-number stragglers (
"$1,234.56","1,200,500.00") in a mostly-plain numeric column are now parsed by the existing locale-aware rescue instead of being coerced to missing — the rescue previously only engaged when the plain parse failed the threshold entirely (gauntlet finding). -
fd.validate_fieldsconsensus inference now honours the same contamination boundary as thefix_dtypeswarning that points users at it (dominant share ≥ 60% with at most a handful of stragglers). Previously the warning fired from a 60% parse share but the consensus gate required 80%, so the exact frame the warning named sailed throughvalidate_fieldssilently (gauntlet finding). -
Explicitly allowed values are no longer swallowed by null-marker heuristics in
fd.validate_fields: withFieldSpec(allowed_values={"US", "DE", "NA"}),"NA"is Namibia, not a missing value (gauntlet finding). -
clean_text/validate_fieldstext normalization no longer rewrites typography in content-bearing fields: forfree_text,textand entity name types, the punctuation→ASCII mapping (curly quotes, em-dashes, prime marks —12″became12") is withheld, matching the field-aware safety contract. Untyped columns keep the existing behaviour (gauntlet finding). -
anonymize()called with norulesand nodetection_confignow emits aUserWarninginstead of silently returning the data unchanged — a privacy call that does nothing must say so. Behavior is otherwise unchanged; pass an empty rule set intentionally by suppressing the warning (found by the installed-wheel matrix audit). -
pip install "freshdata-cleaner[polars]"now actually enables the advertised polars round-trip: the extra was missingpyarrow, whichfd.clean(polars_df)needs for the polars→pandas interchange, so the natural install crashed with polars' internal ModuleNotFoundError. The extra now ships pyarrow, and the adapter raises an actionable message naming the fix when pyarrow is absent (found by the installed-wheel matrix audit). -
explain_cleancell-change reporting: when cleaning removed rows (for example duplicate removal), every shared column previously reported the whole surviving row count as "cells changed". Frames are now aligned on their shared index labels and only genuinely differing cells are counted; cells missing on both sides are unchanged, value↔missing transitions count, and a dtype conversion alone no longer marks untouched values as changed. The elementwise fallback also no longer uses a Python-3.10-onlyzip(strict=...)argument, which crashed on Python 3.9 when reached (#30). -
memory_bytessampled estimation (frames above 200k rows) no longer counts the index payload once per string-like column; a string-heavy index is now measured once, matching the exact path used for smaller frames (#35). -
Integer finalization now checks the exact int64 range in integer space instead of a float magnitude threshold:
-2**63and2**63 - 1024(the largest float64 below2**63) convert to int64/Int64 exactly instead of being demoted to float64, and values at or above2**63can never be admitted by float rounding (#34). -
AI Copilot privacy hardening: sample rows in
model_contextnow hash-mask every string-like column, not only declared/regex-detected PII columns — names, addresses, free text, and obfuscated identifiers in undeclared columns no longer leave the machine raw. A new explicitallow_unmasked_columnsopt-out exists but never exempts declared or detected PII.category_noiseproblem details are stripped of raw value previews before enteringmodel_contextin all privacy modes (includingschema_only); the localreport.problemskeeps the rich previews. -
Out-of-core docs now match measured behavior: keeping a native handle requires
fix_dtypes=Falsein addition tostrategy="conservative"(dtype fixing runs sampled pandas heuristics and forces the recorded fallback), andoutput_format="polars-lazy"defers only the final materialization — pipeline stages currently collect intermediates eagerly, so peak memory during cleaning matches eager output. The DuckDB handle path is the measured lower-peak-memory route (#52, #53). -
CSV formula-injection protection (OWASP):
export_review_queuenow neutralizes spreadsheet formula payloads in CSV exports by default (string cells and column labels starting with= + - @ <tab> <cr>get a leading'; opt out withsanitize_formulas=False) — review queues are built to be opened by humans in spreadsheets.fd.clean_csvand the streaming CLI keep byte-exact output by default and gain an explicit opt-in (sanitize_formulas=True/--sanitize-formulas) covering the cleaned output and the quarantine export. JSONL/Parquet are never altered. -
SECURITY.mdsupported-versions table updated to the current 1.1.x line. -
Source distribution now contains exactly the documented file set: the hatchling
includepatterns are anchored to the repo root, so unanchored names no longer pull in stray matches at any depth (docs/examples/*.html, nestedREADME.mds). -
freshdata-benchmarks/README.mdno longer claims CI execution or published results the repository never produced; it now documents the ASV suite as a locally-run comparative benchmark, separate from the CleanBench CI workflow.
- README rendering on PyPI: the logo and several links (
LICENSE,CHANGELOG.md,CONTRIBUTING.md,CODE_OF_CONDUCT.md,examples/*) used paths relative to the repository, which resolve on GitHub but not on the PyPI project page (rendered with no repository context). All now point to absolutegithub.com/raw.githubusercontent.comURLs. No code changes.
- Interactive output layer (
freshdata.render, lazy-imported):to_html()/_repr_html_()/.show()onCleanReport(collapsible action timeline + filterable audit ledger),Profile(inline quality cockpit),CleanPlan(decision cards / strategy diff grid),ExplainReport(before/after diff explorer), and thecompare_plans/compare_clean/infer_rolesframes (via a transparentReportFrameDataFrame subclass). Self-contained HTML needs no optional deps; new[viz]/[notebook]extras (itables, plotly, great-tables, anywidget) only upgrade the output.Actiongainsstatus/reversible/memory_influenced/human_reviewmetadata. - Cleaning memory:
fd.learn_cleaning_memory/fd.load_cleaning_memory/CleaningMemory(JSON + server-free SQLite storage,to_dict/to_json/diff/summary) andfd.clean(df, memory=...)replay — applies accepted decisions when the dataset signature matches and blocks + explains unsafe replay when the data drifts too far. - Baseline drift convenience:
fd.compare_to_baselinenow accepts a raw DataFrame baseline pluskey=/event_time=for key-level change counts;DriftReportgainswhat_likely_matters()and an interactive view. - Quality-debt ledger:
fd.evaluate_quality_debtscores nine debt dimensions, persists history to SQLite, and escalates warn→fail on repeated/worsening issues. - Dirty-join assistant:
fd.suggest_join_keysproposes exact + fuzzy join keys with confidence, blocking, per-field explanations, and an ambiguous/review section — never auto-joining low-confidence matches. - Text/encoding lint:
fd.lint_text_encodingdetects mixed scripts, mojibake, NFC/NFD inconsistency, RTL/LTR risk, locale-ambiguous dates/numbers, and replacement/control characters (diagnostic-only, with safe-repair flags). - Stakeholder summaries:
fd.stakeholder_summaryexports business-language Markdown / HTML. - Honest out-of-core handles: new
output_format="duckdb"/"polars-lazy"return an un-materialized DuckDB relation / PolarsLazyFrame;CleanReport.materializedflags it. Streaming Polars dedup is now streaming-safe (no forcedmaintain_order) and discloses the order trade-off. - Benchmarks:
benchmarks/bench_report.py(100MB CSV ingest, 1M-row profile, 10M-row null-fill, import-time, memory; balanced vs aggressive) with reproducible commands and honest "not yet measured" placeholders in the docs. - New CDC / event-time quality gate
fd.cdc_profile(df, event_time=..., key=...)(modulefreshdata.cdc, also exportingCDCReport/CDCDefect): classifies change-data-capture defects that are not nulls — stale, late (past-watermark), out-of-order, duplicate-key, invalid-operation, missing-event-time, and replay-risk batches — with per-key ordering, an explicit-watermark mode, and freshness/ordering/CDC trust penalties (each0..1). Read-only; never imputes.CDCReportsupports.summary()/.to_dict()/.to_json()/.to_frame()/.passed/.trust_penalties/.freshness_seconds. - New provenance-aware cleaning for document/OCR-extracted tables (module
freshdata.provenance):fd.clean(df, source_provenance=..., return_report=True)andclean_enterprise(..., source_provenance=...)preserve per-columnsource_file/page/region/parser_confidence/extracted_atand warn when a low-confidence field is coerced or repaired (provenance_confidence_threshold, default0.7). The summary lands atCleanReport.source_provenanceand in.to_dict(). FreshData is the post-extraction normalization/audit layer, not a PDF parser. - New baseline-free contract schema diff (
fd.diff_schema(df, contract=...), exposed lazily fromfreshdata.enterprise.contracts): explains structural schema drift before any repair runs, with no persisted baseline required. Reports added/unexpected, removed, renamed, dtype, nullability, and semantic-domain drift, returning aDriftReportwith a structuredcontract_resultscategorization and.summary()/.to_dict()/.to_json()/.to_frame()exports. Policieson_unexpected(fail|warn|preserve) andon_missing(fail|warn|ignore) control the gate. Rename detection is evidence-based (matching semantic type or high name similarity over a dtype-compatible pair), so unrelated same-dtype columns are never reported as renames.fd.profile(df, contract=...)attaches the same diff atprofile.schema_diff.DriftReportalso gains a.to_frame()exporter shared withmonitor_contract/compare_to_baseline. Read-only; never mutates input. - Contract gate in
fd.cleanandfd.suggest_plan(contract=,on_unexpected=,on_missing=): runsdiff_schemaon the input before repair. A failing gate (errors in the diff) raisesContractViolation(carrying theDriftReportat.report); otherwise the diff is attached to theCleanReportas a JSON-friendlycontract_violationssection that surfaces in.summary()and.to_dict().fd.suggest_plan(df, contract=...)exposes the same diff atplan.schema_diff. In-memory pandas engine only; never auto-renames or drops on the basis of a diff.CleanReportgains acontract_violationsfield. - New wide-schema / large-frame perf controls on
fd.profile:profile_sample=Nprofiles a deterministic N-row sample (stats become estimates),max_columns=Mcaps profiling to the first M columns, andlazy_report=Trueskips the expensive full-frame duplicate-row scan. When any is used theProfiledescribes the profiled subset and records the totals atprofile.materialization(also in.to_dict()).build_profilegains matchingsample=/max_columns=/lazy=keyword-only parameters; default behaviour is unchanged. - New two-frame entity-resolution wrapper
fd.link(left, right, keys=..., strategy="exact"|"fuzzy"|"external")(alsofreshdata.enterprise.link): the ergonomic front door overlink_entities. Builds the resolution config fromkeys+strategy, returns anEntityResolutionReportwith candidate pairs, confidence scores, per-field explanations, and a steward-reviewable structure.strategy="external"formats an adapter callable's pairs (e.g. Dedupe) without re-implementing it. Defaults to the pandas backend (no optional deps); supports ablocking=override andreturn_linked=. - Privacy/regulated-pipeline hardening on
MaskingRule:strategy="token"is now accepted as an alias for the reversibletokenizestrategy, and rules gainretention_days,policy_id, andpolicy_reasonfields.MaskReport(frommask_dataframe) now surfaces per-columnretentionand an auditablepolicy_provenancelist (which rule masked each column, with what strategy, under which policy id, and why), both exported via.to_dict(). FreshData records the declared retention policy for audit; it does not enforce deletion and makes no automatic compliance claims. - New compliance-grade privacy policy engine (
freshdata.enterprise.privacy_policy, exposed asfd.PrivacyPolicy/fd.PrivacyRule/fd.CompliancePack/fd.Jurisdiction/fd.apply_privacy_policy/fd.load_privacy_policy/fd.load_compliance_pack): turns the masking primitives into a declarative, jurisdiction-aware (US / EU / UK / India / Global) policy with actionsclassify/tokenize/pseudonymize/redact/drop/minimize/quarantine/preserve_with_reason. Ships built-in HIPAA, FERPA, PCI and GDPR rule packs (YAML underfreshdata/compliance/packs) combining column-name, value-regex, context and entity/domain-pack classifiers; PCI card numbers are gated by a Luhn check. Policies load from YAML/JSON. Reversible tokenisation uses pluggable vault backends (memory/json/sqlite, viafd.make_vault) and requires an explicit vault and key;detokenize_seriesreverses only with both. The returnedPrivacyReportgains a Data-Trust privacy dimension (sensitive_fields_detected/_touched,unprotected_sensitive_fields,policy_violations, 0–100 score), per-column audit fields (rule_id,action,legal_basis_or_reason,jurisdiction,compliance_pack), plusto_frame()/to_json(). Reports redact previews and never expose vault secrets by default. The legacydetect_pii/anonymize/check_k_anonymity/MaskingRule/PrivacyReportAPI is unchanged. - New schema-drift & data-contract monitoring (
freshdata.enterprise.contracts, exposed asfd.build_baseline/fd.save_baseline/fd.load_baseline/fd.compare_to_baseline/fd.monitor_contract): record a versioned, PII-safeDatasetBaseline(schema + numeric/categorical/datetime statistics) for a trusted dataset, persist it as JSON ("schema_version": "freshdata-baseline-v1"), then detect schema drift, distribution drift (dependency-free KS statistic and PSI over baseline quantile/frequency bins),DataContractviolations (dtype/nullable/unique/ allowed-values/min-max/regex/cardinality), and a trust-score quality gate. Baselines never store raw sample values unlessinclude_samples=True; category labels are hashed by default. Configured viaDriftConfig. Findings are JSON-serialisable and the input frame is never mutated. - New stronger PII detection + reversible / format-preserving anonymization
(
freshdata.enterprise.privacy, exposed asfd.detect_pii/fd.anonymize/fd.check_k_anonymity): a Presidio-style but dependency-free detector (regex + context keywords, optional Presidio NER behind the[privacy]extra) across 15+ entity types with HIPAA/GDPR context boosting; reversible tokenization with an in-memory or JSONTokenVault(tokenize_value/detokenize_value); surrogate/fpeformat-preserving anonymization (clearly flagged as not cryptographic FPE unlesspyffxis installed); HIPAA/GDPR-taggedMaskingEventaudit records that redact raw previews by default (audit_include_pii=Trueto include them); and acheck_k_anonymityre-identification report.MaskingRulegainstokenize/fpe/surrogatestrategies plusentity_types/reversible/key/key_env/token_vault_path/preserve_format/hipaa_tags/gdpr_tags; all existing strategies keep working unchanged. - New probabilistic entity resolution at scale (
freshdata.enterprise.entity_resolution, exposed asfd.resolve_entities/fd.link_entities): a Splink-style, DuckDB-backed record-linkage backend (with a pandas fallback) that blocks candidate pairs via SQL predicates, scores them with weighted comparisons (exact / Jaro–Winkler / Levenshtein / numeric & date distance / phonetic Soundex / custom SQL — all pure-Python primitives), and builds entity clusters via connected components with a completeness-based canonical record. A hardmax_pairsgate prevents cartesian explosions. Configured viaEntityResolutionConfig/BlockingRule/ComparisonLevel. Documented as rule-weighted probabilistic linkage (not full EM-trained Splink parity). EnterpriseConfiggainsdrift/privacy/anonymization/k_anonymity/entity_resolutionsub-configs andenable_contracts/enable_privacy_detection/enable_entity_resolutiontoggles;clean_enterpriseacceptsbaseline=/contract=andEnterpriseResultnow carriesdrift_report/privacy_report/k_anonymity_report/entity_resolution_report. New optional extras[privacy]and[entity-resolution], plus examplesschema_drift_monitoring.py,privacy_anonymization.py, andentity_resolution_duckdb.py.- New FHIR R4 JSON parser (
fd.parse_domain(source, format="fhir")): flattens a Bundle, a single resource, a list of resources, a JSON string, or a file path intopatient/observation/encounter/condition/medication_requestframes whose columns line up with the healthcare validators. The healthcare pack now validates Condition and MedicationRequest (FHIR R4 clinical-status / status / intent value sets, ICD-10 codes against a documented common sample, ISO-8601 dates), adds UCUM unit validation on Observations via the reference layer, and auto-detects all five resources. Resource IDs are never imputed;patient_idstays PHI-masked unlessaudit_include_phi=True; unsupported resource types are recorded as warnings, not dropped. The HL7 v2 parser now also parses theOBRsegment (anorderframe, with eachOBXlinked to its order). - New format parsers (
freshdata.parsers) andfd.parse_domain/fd.clean_domain_file: structural readers that turn HL7 v2 ER7 (MSH/PID/PV1/OBX → patient/encounter/observation, with LOINC/SNOMED/ICD-10 code-system URIs), GPX (waypoints/routes/tracks), SDMX-ML (audit-only observations), and UN/EDIFACT (segments/elements, honoringUNAdelimiters + the release character) into DataFrames. Parsers register via afreshdata.parsersplugin registry; malformed input is recorded inParseResult.warningsrather than raising. - New centralized reference-data layer (
freshdata.domains.reference): one cached, normalizer-awareload_reference(...)over the bundled code sets (ISO-4217, ISO-3166, UN/CEFACT units, plus new UCUM and UN/LOCODE samples), each with a_metaversion/disclaimer block. Supports case-sensitive/insensitive matching, synonym coercion, and aninvalid_maskfor validators. - New finance tick mode (
fd.clean(df, domain="finance", finance_mode="tick")): validates market tick/trade data — ISO-8601 non-future timestamps, positive price/size, ISO-4217 currency (via the reference layer), non-crossed quotes (bid <= ask), duplicate-tick detection, and BCBS-239 / SOX-style completeness controls. Symbol and exchange are IDs and are never imputed; the defaultfinance_mode="ledger"is unchanged. - New energy (SCADA / Modbus) domain pack:
fd.clean(df, domain="energy")validates point-level telemetry — one row per(asset_id, register_address, timestamp)reading — against common Modbus/SCADA conventions: the 16-bit register-address range (0–65535), the public Modbus function codes (1, 2, 3, 4, 5, 6, 15, 16), OPC/SCADA point quality (good/bad/uncertain/stale/null, with synonym coercion), engineering units, and non-future ISO-8601 timestamps. Asset IDs are never imputed; bad/stale/uncertain readings and function/register-class mismatches are flagged for audit rather than dropped. Bundled reference data ships with_metaversion/disclaimer notes documenting that these are common public conventions, not exhaustive vendor specifications. The validator is stateless per frame, so it composes with micro-batch streaming. - New
freshdata.streamingsubpackage andfd.StreamingCleanerfor streaming / micro-batch cleaning of datasets larger than memory. It consumes pandas (and, when installed, PyArrowTable/RecordBatchand polarsDataFrame/LazyFrame) batches, keeps bounded running statistics across them — Welford mean/variance, reservoir-sampled medians, Space-Saving top-k categories — and emits the same explainableCleanReportper micro-batch, now carrying astreamingblock withbatch_id, rows seen, and per-batch / rolling / cumulative trust scores plus a schema-drift flag. Imputation runs in a warmup phase (collect stats, defer and audit) then a stable phase (impute from running stats), preserving every leakage-aware safety gate (ID/target/free-text). Optional source connectors (clean_kafka,clean_arrow_flight) sit behind newfreshdata[kafka|flight]extras and raise a clearImportErrorwhen absent. New CLI subcommandsfreshdata stream,stream-kafka, andbenchmark-streamprocess CSV/Parquet batch-by-batch with per-batch + summary reports and a trust-gate exit code, andbenchmarks/bench_streaming.pyproves stable memory across a lazily-generated 100M-row stream.CleanReportserialization stays backward compatible (nostreamingkey for normal in-memory cleans). - New
freshdata.executionsubpackage: a pluggable, out-of-core / Arrow-native execution engine.fd.clean()gains keyword-onlyengine("pandas"|"polars"|"duckdb"|"auto"),output_format("pandas"|"polars"|"arrow"), andengine_config(EngineConfig) arguments — all backward compatible; default callers are unchanged. The Polars backend cleansLazyFrame/Parquet sources with projection/predicate pushdown and streaming collection; the DuckDB backend cleans via staged SQL with spill-to-disk under a configurablememory_limit. Both reproduce the deterministic representation-repair + structural-reduction + full-row-dedup subset natively (identicalCleanReportto the pandas pipeline) and transparently fall back to pandas for the accuracy-first decision engine, dtype heuristics, and opt-in impute/outliers.engine="auto"picks a backend from the source type and row count, andfd.clean("data.parquet")now also reads a file path directly. New optional extras:freshdata[polars|duckdb|pyarrow|outofcore|bench]. - New
freshdata.benchmarksharness (python -m freshdata.benchmarks.run_benchmarks) that generates synthetic Parquet at a target row count without materialising it, then timesfd.cleanacross the pandas/polars/duckdb backends (wall time, peak resident memory, throughput, Data Trust Score). Seesrc/freshdata/benchmarks/RESULTS.mdfor a 10k–10M reference run. - New
freshdata.integrationssubpackage with first-class orchestration hooks for Dagster (freshdata_asset_check,FreshDataResource), Airflow (FreshDataCleanOperator), and dbt (FreshDataDbtTransform, thedbt-gateCLI, and afreshdata_trust_gatemacro). A framework-agnostic core,evaluate_trust_gate(df, ...) -> (DataFrame, TrustGateResult), cleans a frame and gates it on the 0-100 Data Trust Score, reacting to a low score with warn / fail / skip. Each adapter is an opt-in extra (freshdata[dagster|airflow|dbt|integrations]) and imports cleanly without its framework; a compliance bundle is attached to the gate report whenfreshdata.complianceis available. - New
freshdata.compliancesubpackage that maps aCleanReportonto regulatory control frameworks and emits standards-grade audit artifacts viagenerate_compliance_report(report, frameworks=[...]) -> ComplianceBundle. Five frameworks ship:21cfr_11(21 CFR §11.10(e) audit trail),gdpr_30(Article 30 + 17),alcoa_plus(ALCOA+ data integrity),sox_404(transformation controls), andhipaa_safe_harbor(18-identifier coverage). Reports are purely additive and report-only (never mutate the input). Optionaldataframe=recovers column roles/missing ratios viainfer_roles, andenterprise_result=folds in the Data Trust Score, PII-masking events, and clustering lineage.ComplianceConfig.strict_cfr_normalization(defaultFalse) toggles whether lossless normalising rewrites count as obscuring for the 21 CFR gate. - Four new domain validator packs:
healthcare(FHIR/US Core —Patient,Observation,Encounterwithfhir_resource=/auto-detection),education(Ed-Fi),agriculture(ADAPT, with area/yield unit coercion), andmedia(EIDR/DDEX viamedia_type=/auto-detection, with tested EIDR Mod 37,2 and ICPN GS1 mod-10 check digits). Healthcare/education redact PHI in the audit trail as[PHI]unlessaudit_include_phi=True.fd.cleangains optionalfhir_resource,media_type, andaudit_include_phikeyword arguments. - P1 repair-layer primitives for validator bridges, schema drift harmonization, duplicate/replay defense, and human review queues.
- Top-level bridge adapters:
freshdata.from_gx,freshdata.from_dbt_failures,freshdata.from_pandera_errors,freshdata.emit_gx_expectations, andfreshdata.emit_dbt_tests.
- Packaging: the distribution is
freshdata-cleaneragain. A recent commit renamed the project back tofreshdata, a name PyPI rejects as too similar to the existingfresh-dataproject (the exact collision that forced the original rename).pyproject.toml, every in-source install hint, the docs, and the packaging tests now agree onpip install freshdata-cleaner(import name unchanged:import freshdata). - MissForest:
<col>_was_missingindicators are no longer all-False. The indicator was computed after the column had been imputed, so it never marked any row (andmissforest_add_indicators="auto"never fired at all). Indicators now come from the pre-fill missing mask and the pre-computed column context, matching the standard imputation engine. - Outliers: an explicit
outlier_actionis now honored. Under the defaultstrategy="balanced",outlier_action="cap"(and"remove") was silently downgraded to"flag", so capping never happened despite being the documented default — extreme values were returned unchanged. Explicit"cap"/"remove"/"flag"are now applied to every eligible numeric column. - Small frames no longer skip outlier handling. The engine's minimum non-null threshold dropped from 10 to 4 (the floor at which IQR / z-score fences are defined), so outliers in small DataFrames are detected and handled.
- The default
outlier_actionis now"auto"(context-aware: flags underbalanced, caps underaggressive, flags heavy-tailed >15%-outlying columns). The default behavior underbalancedis unchanged (still flags); only the explicit-directive path changed. An explicitcap/removeon a heavy-tailed column now caps / removes and emits a warning instead of silently flagging.
- Single-string config fields no longer split into characters. Passing a
bare string such as
id_columns="sku_num"went throughtuple()and became one entry per character, so ID protection,preserve_columns,duplicate_subsetandextra_sentinelssilently matched nothing. These fields now accept either one name or a sequence of names. - An explicit
outlier_action="cap"/"remove"is honored instead of being downgraded to"flag"understrategy="balanced". This shipped in 1.0.1 and is described in full under [1.1.0].
First stable release. The public API is now considered stable under Semantic Versioning — breaking changes will require a 2.0.
- Promoted the package to Production/Stable (
Development Status :: 5).
- No behavioral changes versus 0.5.0. The stable public surface is
fd.clean,fd.profile,fd.suggest_plan,fd.compare_plans,fd.compare_clean,fd.explain_clean,fd.infer_roles,fd.Cleaner,fd.CleanConfig,fd.CleanReport/fd.Action,fd.Profile, and the lazily importedfreshdata.enterpriselayer. - Install:
pip install freshdata-cleaner; import:import freshdata as fd.
- Documentation site built with MkDocs Material and deployed to GitHub
Pages (https://freshcode-org.github.io/freshdata/): installation,
quickstart, cleaning-engine, profiling, feature overview, benchmarks,
auto-generated API reference (mkdocstrings), FAQ, and contributing guides,
with search, dark/light mode, OpenGraph metadata,
sitemap.xml, androbots.txtfor SEO/AI discoverability. examples/— 8 runnable scripts (missing values, outliers, normalization, profiling, ML pipeline, large datasets, pandas integration, CSV automation) andnotebooks/— 3 reproducible Jupyter walkthroughs.- Packaging governance:
MANIFEST.in,SECURITY.md,RELEASE.md,.pre-commit-config.yaml, a tag-triggered PyPI release workflow (release.yml) using Trusted Publishing, a docs-deploy workflow (docs.yml), and an issue-template chooser config. - Expanded PyPI keywords and classifiers and a
Documentationproject URL for better search ranking and discoverability.
clean_enterprise(df)and the reusableFreshDataEnterprisepipeline: core cleaning → fuzzy value clustering → semantic validation → PII masking, returning anEnterpriseResult(cleaned frame + trust scores + quality report + lineage). Accepts and returns pandas or polars — Polars-native on the hot path when installed, with a vectorized pandas fallback otherwise.- Data Trust Score (
compute_trust_score,TrustScore): a 0–100 score from completeness, validity, uniqueness, and structural consistency, with per-column detail and a JSON/MarkdownQualityReport(build_quality_report). - Value clustering (
merge_clusters,cluster_column): OpenRefine-style fingerprint key-collision and n-gram merging of variants/typos, built from native Polars string expressions (pandas fallback), withmost_frequent/longest/shortest/firstcanonicalisation. - PII masking (
mask_dataframe,MaskingRule): salted SHA-256hash,redact,partial,regex_scrub(built-in email/phone/SSN/credit-card/IP/IBAN patterns), anddrop; null-preserving and frame-type-preserving. - Semantic validation (
SemanticValidator+ReferenceSetValidator/RegexValidator/CallableValidator/APISemanticValidator,run_semantic_validation), including a built-in ISO-3166iso_country_validator. - Lineage (
LineageTracker,schema_of): records who/when/input-schema/output-schema/ rule per step and exports OpenLineage-compatibleSTART/COMPLETERunEvents (schema + column-lineage facets) with no hard dependency on the OpenLineage client. - Optional Cleanlab wrappers (
detect_label_issues,detect_outliers) with a clear install-hint error when cleanlab is absent. - CLI (
freshdata):clean/profile/trustsubcommands reading CSV/Parquet/JSON, emitting JSON quality + OpenLineage reports, with a non-zero exit code on trust-gate failure — suitable as an Airflow/Prefect batch step. Config via JSON/YAML files. - New optional-dependency extras:
pyarrow,semantic,cli,cleanlab, aggregateenterprise, andall. Polars/PyArrow/requests/cleanlab are imported lazily, so plainimport freshdatastays dependency-light.
- Default strategy is now
"balanced"— accuracy-first cleaning that preserves high-missing columns, flags outliers instead of capping, and skips KNN imputation. Usestrategy="aggressive"for v0.2-style scrubbing (KNN, column drops, winsorization). strategy="auto"is deprecated (alias for"aggressive"; emitsDeprecationWarningonce per process).
fd.suggest_plan(df)andfd.compare_plans(df)— dry-run previews of engine model choices per column, with ranked alternatives.- Model selection router (
engine/model_select.py) scoring imputation and outlier actions;Action.model_idrecords the chosen model. - Expanded target/label heuristics (
aqi,*_bucket,score, …) and domain-sensitive outlier preservation (pollutants, prices, latency, …). profile(df, include_plan=True)attaches aCleanPlanatprofile.plan.src/freshdata/py.typedmarker for PEP 561 typing support.- Multi-dataset regression suite (
tests/fixtures/,test_regressions.py,test_realworld.py,test_model_select.py,test_plan.py). - Golden report snapshots (
tests/fixtures/golden/,pytest --update-golden). - Benchmark tests (
test_benchmark.py) andbenchmarks/bench.py --fixtures. - CI enforces ≥93% coverage and treats
freshdatawarnings as errors. - README migration guide for 0.2 → 0.3.
- KNN imputation: collinearity pruning, row-count gate (10k), warning suppression, index alignment on fill.
- Re-cleaning idempotency for outlier flag columns.
fd.compare_clean()— side-by-side quality + efficiency metrics per strategy.- Four new scenario fixtures:
large_panel(3k rows),duplicate_heavy,locale_numbers,mixed_roles. - Performance baselines (
tests/fixtures/perf/baselines.json) with 25% regression gate. @pytest.mark.largeoptional full AQI.csv benchmark (FRESHDATA_AQI_PATH).- Engine perf: one-pass
EngineCache(contexts + correlation matrix), lazy informative-missing checks, sampled skew on large columns. benchmarks/bench.py --comparetable output.
fd.clean(df) now performs real, context-aware automatic cleaning by
default, driven by a rule-based decision engine.
- Decision engine (
strategy="auto", the new default): profiles every column (missing ratio, dtype, skewness, cardinality, inferred role, informative missingness) and applies threshold rules for missing values and outliers. Every action — including deliberately preserving a column — is logged with a rationale, risk level, and confidence score. - Missing-value bands with configurable thresholds
(
missing_threshold_low/medium/high, defaults 0.05/0.30/0.60): contextual mean/median/mode/sentinel/ffill imputation, KNN imputation for correlated numeric features (scikit-learn optional), column drops for high/extreme missingness with logged reasons,<col>_was_missingindicator columns when missingness is informative. - Column-role inference: targets are never modified, IDs are never imputed, free text is never force-filled, datetimes use time-aware fills.
- Outlier engine:
outlier_action="cap"(default) /"remove"/"flag"/None;outlier_method="auto"(z-score for ~normal, IQR for skewed) and"isolation_forest"; heavy-tail protection (flag instead of cap); domain-sensitive columns (fraud/anomaly/risk) keep their extremes. - Duplicate rules:
duplicate_keep="first"/"last"/"drop"/"aggregate",duplicate_thresholddata-quality warning, time-indexed frames protected unlessallow_timeseries_duplicates=True; count and percentage reported. - New
clean()parameters:strategy, the threshold options,outlier_action,preserve_original,return_report,verbose,preserve_columns,target_column,id_columns,advanced_imputation,missing_indicators. - Report upgrades: per-action
rationale/risk/confidence, missing cells before/after, duplicates removed, outliers handled, columns dropped/imputed/preserved,warnings,recommendations, and a compactbrief()used byverbose=True. - Optional extra:
pip install "freshdata-cleaner[ml]"for scikit-learn.
- Default behavior: statistical cleaning now runs by default. Pass
strategy="conservative"for the 0.1.x representation-only behavior; explicitimpute=/outliers=still override the engine. report.to_frame()gainedrationale,risk, andconfidencecolumns.verbose=True(default) prints a one-line summary per clean.
Initial release.
freshdata.clean()— automatic, audited cleaning: column-name normalization, whitespace stripping, sentinel-string normalization, empty row/column pruning, validated dtype inference (numeric incl. currency/thousands separators, datetime, boolean), and exact duplicate removal.- Opt-in steps: imputation (
auto/mean/median/mode), outlier clipping/flagging (IQR or z-score), constant-column dropping, memory optimization (numeric downcasting + category conversion), index reset. freshdata.profile()— read-only profiling whose dtype suggestions are produced by the same inference codecleanuses.freshdata.Cleaner— reusable configured pipeline withreport_.freshdata.CleanConfig— frozen, self-validating configuration; unknown options raise with a "did you mean" suggestion.freshdata.CleanReport/freshdata.Action— structured audit trail withsummary(),to_dict(),to_frame().- Type hints throughout (
py.typed), zero dependencies beyond pandas/numpy, support for Python 3.9–3.13.