Skip to content

Commit a6f5a71

Browse files
Merge pull request #96 from FreshCode-Org/feature/phase4-learning-profiles-jwd
[codex] Align package install and imputation UX
2 parents 635abe5 + cb2ba31 commit a6f5a71

42 files changed

Lines changed: 11345 additions & 93 deletions

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎README.md‎

Lines changed: 25 additions & 18 deletions
Original file line numberDiff line numberDiff line change
@@ -10,12 +10,12 @@
1010
logs, scores, explains, and *remembers* data-quality decisions, then makes them
1111
reusable in notebooks, streaming runs, and orchestrated pipelines.
1212

13-
[![PyPI Version](https://img.shields.io/pypi/v/freshdata-cleaner.svg)](https://pypi.org/project/freshdata-cleaner/)
14-
[![Python Versions](https://img.shields.io/pypi/pyversions/freshdata-cleaner.svg)](https://pypi.org/project/freshdata-cleaner/)
13+
[![PyPI Version](https://img.shields.io/pypi/v/freshdata.svg)](https://pypi.org/project/freshdata/)
14+
[![Python Versions](https://img.shields.io/pypi/pyversions/freshdata.svg)](https://pypi.org/project/freshdata/)
1515
[![License: MIT](https://img.shields.io/badge/license-MIT-green.svg)](LICENSE)
1616
[![CI](https://github.com/FreshCode-Org/freshdata/actions/workflows/ci.yml/badge.svg)](https://github.com/FreshCode-Org/freshdata/actions/workflows/ci.yml)
1717
[![Docs](https://github.com/FreshCode-Org/freshdata/actions/workflows/docs.yml/badge.svg)](https://freshcode-org.github.io/freshdata/)
18-
[![Downloads](https://img.shields.io/pypi/dm/freshdata-cleaner.svg)](https://pypi.org/project/freshdata-cleaner/)
18+
[![Downloads](https://img.shields.io/pypi/dm/freshdata.svg)](https://pypi.org/project/freshdata/)
1919
[![Coverage](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/FreshCode-Org/freshdata/badges/coverage.json)](https://github.com/FreshCode-Org/freshdata/actions/workflows/ci.yml)
2020
[![Ruff](https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/astral-sh/ruff/main/assets/badge/v2.json)](https://github.com/astral-sh/ruff)
2121
[![Checked with mypy](https://img.shields.io/badge/mypy-checked-blue.svg)](https://mypy-lang.org/)
@@ -47,9 +47,12 @@ import freshdata as fd
4747

4848
df = pd.read_csv("export.csv")
4949

50-
cleaned = fd.clean(df) # one line
51-
cleaned, report = fd.clean(df, return_report=True) # ... with a full audit trail
52-
print(report.summary())
50+
result = fd.clean(df) # one line
51+
cleaned = result.data # plain pandas DataFrame
52+
print(result.summary()) # full audit trail
53+
result.visualize() # self-contained HTML
54+
55+
cleaned, report = fd.clean(df, return_report=True) # legacy tuple still works
5356
```
5457

5558
```text
@@ -118,13 +121,13 @@ ML-ready data without writing — or trusting — yet another bespoke script.
118121
## 📦 Installation
119122

120123
```bash
121-
pip install freshdata-cleaner # pandas + numpy only
122-
pip install "freshdata-cleaner[ml]" # + scikit-learn (KNN imputation, IsolationForest)
123-
pip install "freshdata-cleaner[domains]" # + PyYAML (finance, GS1, and GTFS packs)
124-
pip install "freshdata-cleaner[enterprise]" # + polars, pyarrow, requests, pyyaml (enterprise layer + CLI)
125-
pip install "freshdata-cleaner[privacy]" # + Presidio NER & pyffx (stronger PII detection / crypto FPE)
126-
pip install "freshdata-cleaner[entity-resolution]" # + duckdb (probabilistic linkage at scale)
127-
pip install "freshdata-cleaner[all]" # everything, including cleanlab
124+
pip install freshdata # pandas + numpy, reporting, standard HTML visualization
125+
pip install "freshdata[ml]" # + scikit-learn (KNN imputation, IsolationForest)
126+
pip install "freshdata[domains]" # + PyYAML (finance, GS1, and GTFS packs)
127+
pip install "freshdata[enterprise]" # + polars, pyarrow, requests, pyyaml (enterprise layer + CLI)
128+
pip install "freshdata[privacy]" # + Presidio NER & pyffx (stronger PII detection / crypto FPE)
129+
pip install "freshdata[entity-resolution]" # + duckdb (probabilistic linkage at scale)
130+
pip install "freshdata[all]" # everything, including cleanlab
128131
```
129132

130133
Requires **Python ≥ 3.9** and **pandas ≥ 1.5**. Verify the install:
@@ -142,12 +145,15 @@ import freshdata as fd
142145
df = pd.read_csv("messy_export.csv")
143146

144147
# Clean with sensible, explainable defaults
145-
cleaned, report = fd.clean(df, return_report=True)
148+
result = fd.clean(df)
149+
cleaned = result.data
150+
report = result.report()
146151

147152
print(report.summary()) # human-readable audit trail
148153
report.to_frame() # decisions as a DataFrame
149154
report.to_dict() # JSON-friendly for logging / dashboards
150-
report.show() # interactive action timeline + audit ledger (notebook)
155+
result.visualize() # self-contained HTML action timeline + audit ledger
156+
report.show() # inline in notebooks, or writes a standalone .html file
151157
```
152158

153159
Interactive output, decision memory, drift, debt, joins, encoding, and
@@ -170,7 +176,8 @@ lint = fd.lint_text_encoding(df, columns=["name", "city"])
170176
brief = fd.stakeholder_summary(report, audience="business", format="markdown")
171177
```
172178

173-
Visualization extras are optional and never required by the base install:
179+
Standard report visualization is included in the base install. Optional
180+
visualization extras only add richer third-party notebook/table integrations:
174181
`pip install 'freshdata[viz]'` (or `[notebook]`, `[all]`).
175182

176183
Domain packs add versioned validation and separately audited repairs:
@@ -297,7 +304,7 @@ categorical, and boolean predictors, while preserving FreshData's role gates:
297304
targets, IDs, and free-text columns are not fabricated.
298305

299306
```python
300-
# pip install "freshdata-cleaner[ml]"
307+
# pip install "freshdata[ml]"
301308
cleaned, report = fd.clean(
302309
df,
303310
impute_method="missforest",
@@ -410,7 +417,7 @@ It accepts **pandas**, and (when installed) **PyArrow** `Table`/`RecordBatch` an
410417
stream. Optional source connectors live behind extras:
411418

412419
```python
413-
# pip install "freshdata-cleaner[kafka]" / "freshdata-cleaner[flight]"
420+
# pip install "freshdata[kafka]" / "freshdata[flight]"
414421
cleaner.clean_kafka(topic="events", bootstrap_servers="localhost:9092", batch_size=10_000)
415422
cleaner.clean_arrow_flight("grpc://localhost:8815", batch_size=100_000)
416423
```

‎docs/backends.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -84,7 +84,7 @@ cfg = EngineConfig(engine="duckdb", memory_limit_gb=4, temp_directory="/tmp/spil
8484
cfg = EngineConfig(engine="spark", spark_shuffle_partitions=200, output_format="spark")
8585
```
8686

87-
PySpark is an **optional dependency** (`pip install 'freshdata-cleaner[spark]'`) and
87+
PySpark is an **optional dependency** (`pip install 'freshdata[spark]'`) and
8888
also needs a JVM at runtime. Importing `freshdata` never imports pyspark.
8989

9090
FreshCore is also optional. Install the Python package normally, then build the

‎docs/benchmarks.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -27,7 +27,7 @@ gated to aggressive mode only.
2727
MissForest-style imputation is also benchmarked separately because it is
2828
opt-in, scikit-learn-backed, and intentionally slower than the default engine.
2929
Use `python benchmarks/bench_missforest.py` after installing
30-
`freshdata-cleaner[ml]` to compare median/mode, aggressive KNN, and MissForest
30+
`freshdata[ml]` to compare median/mode, aggressive KNN, and MissForest
3131
on mixed-type synthetic data.
3232

3333
## The Benchmark Release harness

‎docs/cleaning-engine.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -77,7 +77,7 @@ mixed tabular data, you can opt into MissForest-style random-forest imputation
7777
with the optional ML extra:
7878

7979
```bash
80-
pip install "freshdata-cleaner[ml]"
80+
pip install "freshdata[ml]"
8181
```
8282

8383
```python
@@ -130,7 +130,7 @@ fd.clean(
130130
target_column="churn",
131131
preserve_columns=("notes",),
132132
id_columns=("ref",),
133-
impute_method="missforest", # optional; requires freshdata-cleaner[ml]
133+
impute_method="missforest", # optional; requires freshdata[ml]
134134
missforest_max_iter=5,
135135
missforest_n_estimators=100,
136136
return_report=True,

‎docs/faq.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -47,7 +47,7 @@ you explicitly pass `preserve_original=False` to reuse memory.
4747

4848
## Does it support Polars?
4949

50-
Yes — install `pip install "freshdata-cleaner[polars]"` and pass a Polars DataFrame to
50+
Yes — install `pip install "freshdata[polars]"` and pass a Polars DataFrame to
5151
`fd.clean`; you get a Polars DataFrame back.
5252

5353
## How do I see what it changed?

‎docs/feature-overview.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -71,4 +71,4 @@ import freshdata as fd
7171
cleaned = fd.clean(pl_df) # returns a pl.DataFrame when the input is Polars
7272
```
7373

74-
Install with `pip install "freshdata-cleaner[polars]"`.
74+
Install with `pip install "freshdata[polars]"`.

‎docs/index.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -43,7 +43,7 @@ score for each decision.
4343
## Install
4444

4545
```bash
46-
pip install freshdata-cleaner
46+
pip install freshdata
4747
```
4848

4949
See [Installation](installation.md) for optional extras (`ml`, `enterprise`, `all`).

‎docs/installation.md‎

Lines changed: 11 additions & 10 deletions
Original file line numberDiff line numberDiff line change
@@ -3,7 +3,7 @@ title: Installation
33
description: >-
44
How to install freshdata for Python and pandas, including optional extras for
55
scikit-learn machine-learning imputation, the enterprise layer, and Polars.
6-
keywords: install freshdata, pip install freshdata-cleaner, pandas data cleaning install
6+
keywords: install freshdata, pip install freshdata, pandas data cleaning install
77
---
88

99
# Installation
@@ -13,11 +13,12 @@ keywords: install freshdata, pip install freshdata-cleaner, pandas data cleaning
1313
## Basic install
1414

1515
```bash
16-
pip install freshdata-cleaner
16+
pip install freshdata
1717
```
1818

19-
This installs the pandas + NumPy core — everything you need for `fd.clean`,
20-
`fd.profile`, and the decision engine.
19+
This installs the pandas + NumPy core plus FreshData's standard reporting and
20+
self-contained HTML visualization. You do not need an extra to call
21+
`fd.clean(df).summary()`, `fd.clean(df).report()`, or `fd.clean(df).visualize()`.
2122

2223
## Optional extras
2324

@@ -26,7 +27,7 @@ Install only what you need:
2627
=== "Machine learning"
2728

2829
```bash
29-
pip install "freshdata-cleaner[ml]"
30+
pip install "freshdata[ml]"
3031
```
3132

3233
Adds **scikit-learn** for KNN imputation and IsolationForest outlier
@@ -35,7 +36,7 @@ Install only what you need:
3536
=== "Enterprise"
3637

3738
```bash
38-
pip install "freshdata-cleaner[enterprise]"
39+
pip install "freshdata[enterprise]"
3940
```
4041

4142
Adds **polars, pyarrow, requests, pyyaml** for the enterprise layer:
@@ -45,15 +46,15 @@ Install only what you need:
4546
=== "Everything"
4647

4748
```bash
48-
pip install "freshdata-cleaner[all]"
49+
pip install "freshdata[all]"
4950
```
5051

5152
All extras above plus **cleanlab** for ML label-noise detection.
5253

5354
=== "Polars only"
5455

5556
```bash
56-
pip install "freshdata-cleaner[polars]"
57+
pip install "freshdata[polars]"
5758
```
5859

5960
Pass a Polars DataFrame to `fd.clean` and get a Polars DataFrame back.
@@ -74,11 +75,11 @@ print(fd.clean(df))
7475

7576
## Note on naming
7677

77-
The PyPI distribution is **`freshdata-cleaner`**, but the import name is simply
78+
The PyPI distribution is **`freshdata`**, but the import name is simply
7879
**`freshdata`** — so you install one and import the other:
7980

8081
```bash
81-
pip install freshdata-cleaner
82+
pip install freshdata
8283
```
8384

8485
```python

‎docs/interactive.md‎

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -18,8 +18,11 @@ visualization library.
1818
```python
1919
import freshdata as fd
2020

21-
cleaned, report = fd.clean(df, return_report=True)
21+
result = fd.clean(df)
22+
cleaned = result.data
23+
report = result.report()
2224

25+
result.visualize() # self-contained report HTML string
2326
report.show() # collapsible action timeline + filterable audit ledger
2427
fd.profile(df).show() # inline quality cockpit
2528
fd.suggest_plan(df).show() # per-column decision cards

‎docs/limitations.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -94,7 +94,7 @@ and cleaning memory need **none** of these.
9494
## Optional semantic models (Phase 3)
9595

9696
- **The default install stays model-free**; everything below applies only
97-
after `pip install "freshdata-cleaner[semantic]"` *and* an explicit
97+
after `pip install "freshdata[semantic]"` *and* an explicit
9898
`fd.models.pull(...)` (or air-gapped file placement). Nothing is ever
9999
downloaded during cleaning.
100100
- **Official model artifacts are not hosted yet.** `fd.models.pull` raises a

0 commit comments

Comments
 (0)