1010logs, scores, explains, and * remembers* data-quality decisions, then makes them
1111reusable in notebooks, streaming runs, and orchestrated pipelines.
1212
13- [ ![ PyPI Version] ( https://img.shields.io/pypi/v/freshdata-cleaner .svg )] ( https://pypi.org/project/freshdata-cleaner / )
14- [ ![ Python Versions] ( https://img.shields.io/pypi/pyversions/freshdata-cleaner .svg )] ( https://pypi.org/project/freshdata-cleaner / )
13+ [ ![ PyPI Version] ( https://img.shields.io/pypi/v/freshdata.svg )] ( https://pypi.org/project/freshdata/ )
14+ [ ![ Python Versions] ( https://img.shields.io/pypi/pyversions/freshdata.svg )] ( https://pypi.org/project/freshdata/ )
1515[ ![ License: MIT] ( https://img.shields.io/badge/license-MIT-green.svg )] ( LICENSE )
1616[ ![ CI] ( https://github.com/FreshCode-Org/freshdata/actions/workflows/ci.yml/badge.svg )] ( https://github.com/FreshCode-Org/freshdata/actions/workflows/ci.yml )
1717[ ![ Docs] ( https://github.com/FreshCode-Org/freshdata/actions/workflows/docs.yml/badge.svg )] ( https://freshcode-org.github.io/freshdata/ )
18- [ ![ Downloads] ( https://img.shields.io/pypi/dm/freshdata-cleaner .svg )] ( https://pypi.org/project/freshdata-cleaner / )
18+ [ ![ Downloads] ( https://img.shields.io/pypi/dm/freshdata.svg )] ( https://pypi.org/project/freshdata/ )
1919[ ![ Coverage] ( https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/FreshCode-Org/freshdata/badges/coverage.json )] ( https://github.com/FreshCode-Org/freshdata/actions/workflows/ci.yml )
2020[ ![ Ruff] ( https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json )] ( https://github.com/astral-sh/ruff )
2121[ ![ Checked with mypy] ( https://img.shields.io/badge/mypy-checked-blue.svg )] ( https://mypy-lang.org/ )
@@ -47,9 +47,12 @@ import freshdata as fd
4747
4848df = pd.read_csv(" export.csv" )
4949
50- cleaned = fd.clean(df) # one line
51- cleaned, report = fd.clean(df, return_report = True ) # ... with a full audit trail
52- print (report.summary())
50+ result = fd.clean(df) # one line
51+ cleaned = result.data # plain pandas DataFrame
52+ print (result.summary()) # full audit trail
53+ result.visualize() # self-contained HTML
54+
55+ cleaned, report = fd.clean(df, return_report = True ) # legacy tuple still works
5356```
5457
5558``` text
@@ -118,13 +121,13 @@ ML-ready data without writing — or trusting — yet another bespoke script.
118121## 📦 Installation
119122
120123``` bash
121- pip install freshdata-cleaner # pandas + numpy only
122- pip install " freshdata-cleaner [ml]" # + scikit-learn (KNN imputation, IsolationForest)
123- pip install " freshdata-cleaner [domains]" # + PyYAML (finance, GS1, and GTFS packs)
124- pip install " freshdata-cleaner [enterprise]" # + polars, pyarrow, requests, pyyaml (enterprise layer + CLI)
125- pip install " freshdata-cleaner [privacy]" # + Presidio NER & pyffx (stronger PII detection / crypto FPE)
126- pip install " freshdata-cleaner [entity-resolution]" # + duckdb (probabilistic linkage at scale)
127- pip install " freshdata-cleaner [all]" # everything, including cleanlab
124+ pip install freshdata # pandas + numpy, reporting, standard HTML visualization
125+ pip install " freshdata[ml]" # + scikit-learn (KNN imputation, IsolationForest)
126+ pip install " freshdata[domains]" # + PyYAML (finance, GS1, and GTFS packs)
127+ pip install " freshdata[enterprise]" # + polars, pyarrow, requests, pyyaml (enterprise layer + CLI)
128+ pip install " freshdata[privacy]" # + Presidio NER & pyffx (stronger PII detection / crypto FPE)
129+ pip install " freshdata[entity-resolution]" # + duckdb (probabilistic linkage at scale)
130+ pip install " freshdata[all]" # everything, including cleanlab
128131```
129132
130133Requires ** Python ≥ 3.9** and ** pandas ≥ 1.5** . Verify the install:
@@ -142,12 +145,15 @@ import freshdata as fd
142145df = pd.read_csv(" messy_export.csv" )
143146
144147# Clean with sensible, explainable defaults
145- cleaned, report = fd.clean(df, return_report = True )
148+ result = fd.clean(df)
149+ cleaned = result.data
150+ report = result.report()
146151
147152print (report.summary()) # human-readable audit trail
148153report.to_frame() # decisions as a DataFrame
149154report.to_dict() # JSON-friendly for logging / dashboards
150- report.show() # interactive action timeline + audit ledger (notebook)
155+ result.visualize() # self-contained HTML action timeline + audit ledger
156+ report.show() # inline in notebooks, or writes a standalone .html file
151157```
152158
153159Interactive output, decision memory, drift, debt, joins, encoding, and
@@ -170,7 +176,8 @@ lint = fd.lint_text_encoding(df, columns=["name", "city"])
170176brief = fd.stakeholder_summary(report, audience = " business" , format = " markdown" )
171177```
172178
173- Visualization extras are optional and never required by the base install:
179+ Standard report visualization is included in the base install. Optional
180+ visualization extras only add richer third-party notebook/table integrations:
174181` pip install 'freshdata[viz]' ` (or ` [notebook] ` , ` [all] ` ).
175182
176183Domain packs add versioned validation and separately audited repairs:
@@ -297,7 +304,7 @@ categorical, and boolean predictors, while preserving FreshData's role gates:
297304targets, IDs, and free-text columns are not fabricated.
298305
299306``` python
300- # pip install "freshdata-cleaner [ml]"
307+ # pip install "freshdata[ml]"
301308cleaned, report = fd.clean(
302309 df,
303310 impute_method = " missforest" ,
@@ -410,7 +417,7 @@ It accepts **pandas**, and (when installed) **PyArrow** `Table`/`RecordBatch` an
410417stream. Optional source connectors live behind extras:
411418
412419``` python
413- # pip install "freshdata-cleaner [kafka]" / "freshdata-cleaner [flight]"
420+ # pip install "freshdata[kafka]" / "freshdata[flight]"
414421cleaner.clean_kafka(topic = " events" , bootstrap_servers = " localhost:9092" , batch_size = 10_000 )
415422cleaner.clean_arrow_flight(" grpc://localhost:8815" , batch_size = 100_000 )
416423```
0 commit comments