Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
8fc67cf
fix(parsers): detect XML encoding and reject DTDs with an expat pass
kevincostner17 Sep 15, 2026
081206d
fix(csv): guard every header level, index labels and axis names again…
kevincostner17 Sep 15, 2026
b3a5a34
fix(duckdb): spill into a private per-run directory instead of /tmp/f…
kevincostner17 Sep 15, 2026
6bc6b83
fix(contracts): keyed or label-free baseline category labels
kevincostner17 Sep 15, 2026
092d233
fix(privacy): create token vault files owner-only
kevincostner17 Sep 15, 2026
cd5427f
fix(privacy): random per-call key for keyless tokenize, surrogate, fp…
kevincostner17 Sep 15, 2026
1bf9e36
fix(learning): mask every detected PII type in privacy='mask' profiles
kevincostner17 Sep 15, 2026
6a74471
fix(enterprise): keep masked column values out of clean_enterprise re…
kevincostner17 Sep 15, 2026
13eaeda
fix(report): key the tokens that stand in for sensitive values
kevincostner17 Sep 15, 2026
53eaa9e
fix(copilot): mask every non-numeric sample column before model_context
kevincostner17 Sep 15, 2026
0ede986
fix(copilot): mask sample values by column position
kevincostner17 Sep 15, 2026
ac9a1d3
feat(copilot): add mask_salt for a reproducible model_context
kevincostner17 Sep 15, 2026
cb9080e
docs(copilot): document positional masking and mask_salt
kevincostner17 Sep 15, 2026
640bad2
test(truthbench): pin copilot canary scanning and sink determinism
kevincostner17 Sep 15, 2026
50efe91
docs: changelog for the security fixes
kevincostner17 Sep 15, 2026
4b3cd9a
fixup(enterprise): don't apply strict column matching to report redac…
kevincostner17 Sep 15, 2026
25f5a59
chore(release): 2.1.0
kevincostner17 Sep 15, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 0 additions & 1 deletion .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,6 @@ site/
# Out-of-core engine: benchmark outputs, synthetic data, and DuckDB spill
src/freshdata/benchmarks/results/
/tmp/freshdata_bench/
/tmp/freshdata_spill/

# Benchmark runtime results: raw case files stay local; compact evidence is committed.
benchmarks/results/*
Expand Down
54 changes: 53 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,53 @@ All notable changes to this project are documented here. The format follows
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and the project
adheres to [Semantic Versioning](https://semver.org/).

## [Unreleased]
## [2.1.0] - 2026-09-15

### Security
- GPX and SDMX parsing detects the document encoding (BOM, UTF-16/UTF-32
prefixes, XML declaration) and checks with expat before reading, so DTD and
entity declarations are rejected in every encoding. Previously a UTF-16
document bypassed the check and allowed entity expansion.
- CSV formula sanitising now covers every level of multi-row headers, and index
labels and names, so crafted header cells from the input are no longer written
as live formulas.
- The DuckDB engine no longer spills to the shared `/tmp/freshdata_spill`.
`EngineConfig.temp_directory` defaults to `None`, and each run spills into a
private (0700), per-run directory under the user's cache directory (or
`FRESHDATA_SPILL_DIR`), which is removed afterwards. An explicit
`temp_directory` is checked for ownership and permissions.
- Baseline category labels are no longer unkeyed SHA-1. Without `label_key` a
baseline stores a label-free frequency profile; with `label_key` (or
`FRESHDATA_BASELINE_KEY`) labels are HMAC-SHA256. Baselines are written as
schema `freshdata-baseline-v2`; v1 baselines still load with a warning and
should be rebuilt.
- `JsonTokenVault` and `SqliteTokenVault` create their files owner-only (0600)
at creation time. SQLite journal/WAL files inherit that mode. Existing
group/other-readable vault files trigger a warning.
- `tokenize`, `surrogate` and keyless `fpe` masking rules, and the policy
`pseudonymize` action (the default in the GDPR, HIPAA and FERPA packs), no
longer fall back to public constants when no key is set. They use a random
per-call key and emit `EphemeralKeyWarning`; pass `key=`/`key_env=` for
stable, joinable output.
- `fd.learn(privacy='mask')` treats every PII type `detect_pii` reports (payment
cards, IBANs, IP addresses, health and licence identifiers) as sensitive;
unknown types fail closed. It adds card and bank column-name hints.
`freshdata profile audit` flags raw card numbers and IBANs in existing
profiles.
- `clean_enterprise` reports no longer contain raw values of masked columns:
cluster canonical, variant and key values, semantic-validation invalid
samples, and `clean_report.coerced_cells` originals (and the coercion warnings
quoting them) for masked columns are masked or redacted. The
`[SENSITIVE:xxxxxxxx]` tokens that stand in for declared `sensitive_columns`
values are now a truncated HMAC-SHA256 under a random per-process key instead
of an unkeyed SHA-256; they still match within a run but differ between runs.
- Copilot sample masking uses an allow-list: only numeric and boolean sample
values pass through, so Arrow-backed string, dictionary and other non-numeric
columns (including datetimes) are hash-masked.
- Copilot masks sample values by column position, so integer, float and tuple
column labels no longer bypass masking. `sensitive_columns` and `must_mask`
match non-string labels, unknown `sensitive_columns` raise, labels that
collide as strings raise, and masking fails closed.

### Added
- `fd.clean_excel()`, the Excel companion to `fd.clean_csv()`: reads one sheet,
Expand Down Expand Up @@ -509,6 +555,12 @@ adheres to [Semantic Versioning](https://semver.org/).
- The `quarantine` privacy-policy action works on nullable integer, boolean
and categorical columns instead of raising `TypeError`; those columns come
back as object dtype and missing cells stay missing.
- `detect_pii` and detection-driven `anonymize` scan categorical text columns on
pandas 1.5, as on pandas 2 (#280).
- Crypto FPE honours `visible` (#281).
- `analyze_dataset(mask_salt=...)` makes `model_context` and its fingerprint
reproducible; by default they are per-run, and `audit["mask_salt_source"]`
records which was used (#288).

## [2.0.0] - 2026-07-20

Expand Down
9 changes: 8 additions & 1 deletion benchmarks/truthbench/surfaces/copilot.py
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,10 @@
from ..privacy import SinkScanner
from .base import ExceptionDetails, SurfaceAdapter, SurfaceObservation, register_adapter

#: Pinned so masked sample tokens, and so every rendered sink, are identical
#: across runs and repeats instead of depending on a per-run random key.
COPILOT_MASK_SALT = "truthbench-fixed-copilot-mask-salt"


class CopilotAdapter(SurfaceAdapter):
"""Run the deterministic Copilot path and retain all report sinks safely."""
Expand All @@ -35,7 +39,10 @@ def observe(self, fixture: Any, context: Any) -> SurfaceObservation:
)
with contextlib.redirect_stdout(stdout), contextlib.redirect_stderr(stderr):
report = analyze_dataset(
frame, provider=None, sensitive_columns=sensitive
frame,
provider=None,
sensitive_columns=sensitive,
mask_salt=COPILOT_MASK_SALT,
)
# The prompt is constructed exactly as a provider call would see it,
# even though this adapter intentionally supplies no provider.
Expand Down
53 changes: 41 additions & 12 deletions docs/ai-copilot.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,11 +41,15 @@ Three properties make this different from "ask a chatbot about my data":
- **Deterministic and offline.** The analysis is rule-based, built from
freshdata's own primitives (profiling, PII detection, the context-policy
compiler, value clustering, trust scoring). The same input always produces
the same report; it runs in CI with no API key and no network access.
the same findings, plan and code; it runs in CI with no API key and no
network access. Masked sample tokens use a per-run key unless you pass
`mask_salt`, so `model_context` and its fingerprint are reproducible only
with a pinned salt.
- **Privacy-first.** Raw string values never enter `report.model_context` —
the only payload an LLM provider would ever see. Every string-like sample
column is hash-masked first (numeric values pass through as-is), or samples
are omitted entirely with `privacy="schema_only"`.
the only payload an LLM provider would ever see. Every sample column that
is not numeric or boolean is hash-masked first (numeric and boolean values
pass through as-is), or samples are omitted entirely with
`privacy="schema_only"`.
- **Actionable.** The output is not advice — it is an ordered plan with a
rationale per step, plus a generated freshdata pipeline you can run as-is.
(The test suite literally `exec()`s the generated code and asserts the
Expand Down Expand Up @@ -97,14 +101,28 @@ artifact you would get from `fd.compile_context`. Unknown rules raise a
The `privacy` parameter controls what goes into `report.model_context`:

- `"mask_pii_before_reasoning"` (default) — includes `sample_rows` sample
rows, but every string-like column is hash-masked first: `must_mask`
columns, columns the PII detector flagged, **and** every other
object/string/categorical column — regex detection cannot see names,
addresses, or free text, so no string value is sent raw. Numeric values
pass through as-is; numeric quasi-identifiers (e.g. exact salary + age)
are the residual risk — drop such columns first or use `"schema_only"`.
rows, but only numeric and boolean columns pass through raw (an
allow-list). Everything else is hash-masked first: `must_mask` columns,
columns the PII detector flagged, **and** every other column —
object/string, Arrow-backed string or dictionary (e.g. from
`read_csv(dtype_backend="pyarrow")`), categorical, bytes, datetime,
timedelta, period, and any dtype the copilot does not recognise. Regex
detection cannot see names, addresses, free text, or dates of birth, so
no such value is sent raw. Numeric quasi-identifiers (e.g. exact salary +
age) are the residual risk — drop such columns first or use
`"schema_only"`.
`allow_unmasked_columns=[...]` is an explicit per-column opt-out; it never
exempts a declared or detected PII column.
exempts a declared or detected PII column. `sensitive_columns=[...]`
declares columns that are always masked, whatever their dtype (an SSN
stored as an integer, an internal case ID).
Column names in `context_policy`, `sensitive_columns` and
`allow_unmasked_columns` match a column's label or its `str()` form, so
integer, float and tuple labels (e.g. from `read_csv(header=None)`) work;
unknown `sensitive_columns` / `allow_unmasked_columns` names raise, and so
do labels that collide once converted to `str` (`0` and `"0"`). Masking is
done by column position and fails closed: if a selected column does not
come back hash-masked, `analyze_dataset` raises `RuntimeError` instead of
building `model_context`.
- `"schema_only"` — no cell values at all; only column names, dtypes,
missing percentages, and aggregate statistics.

Expand All @@ -123,6 +141,16 @@ Two details worth knowing:
- `report.audit["model_context_sha256"]` fingerprints the exact payload a
provider would have seen, so you can prove after the fact what was (and
was not) shared.
- Masked sample tokens are HMAC-SHA256 hashes with a separate salt per
column, so equal values in two columns get different tokens. By default
the salts come from a random per-run key: the same frame gives different
tokens and a different `model_context_sha256` on every run. Pass
`mask_salt="..."` to derive the salts from your value instead, which makes
`model_context` and its fingerprint reproducible (useful in CI). Treat
that value as a secret, since anyone holding it can confirm guesses of
low-cardinality values; it is never written to the report, and
`report.audit["mask_salt_source"]` records only `"caller"` or
`"per-run-random"`.

## Plugging in an LLM (optional, experimental)

Expand Down Expand Up @@ -178,7 +206,8 @@ every time. What the copilot (and freshdata underneath it) adds:
which spellings are the same category — with severity and evidence;
- an audit trail a reviewer can read (`CleanReport` actions with rationale,
masked-context SHA, privacy events with HIPAA/GDPR tags);
- reproducibility: the same input produces the same report, plan, and code.
- reproducibility: the same input produces the same findings, plan, and
code, and with `mask_salt` the same `model_context` and fingerprint.

## Limitations and responsible use

Expand Down
16 changes: 15 additions & 1 deletion docs/backends.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,10 +93,24 @@ interpolation) can differ from the pandas reference.
```python
from freshdata.execution import EngineConfig

cfg = EngineConfig(engine="duckdb", memory_limit_gb=4, temp_directory="/tmp/spill")
import os

cfg = EngineConfig(engine="duckdb", memory_limit_gb=4,
temp_directory=os.path.expanduser("~/scratch/freshdata-spill"))
cfg = EngineConfig(engine="spark", spark_shuffle_partitions=200, output_format="spark")
```

DuckDB spill files contain rows of the data being cleaned, so each run spills into
its own private (0700) subdirectory, removed when the run's connection closes (for
`output_format="duckdb"`, when the returned relation is released). By default the
subdirectory is created under `$FRESHDATA_SPILL_DIR`, or under the per-user cache
directory (`~/.cache/freshdata/spill` or `$XDG_CACHE_HOME/freshdata/spill` on Linux,
`~/Library/Caches/freshdata/spill` on macOS, `%LOCALAPPDATA%\freshdata\spill` on
Windows), falling back to the system temp directory only when that is not writable.
An explicit `temp_directory` is created with mode 0700 if missing; one that is not
owned by you, or is group/other-writable without the sticky bit, raises
`PermissionError`.

PySpark is an **optional dependency** (`pip install 'freshdata-cleaner[spark]'`) and
also needs a JVM at runtime. Importing `freshdata` never imports pyspark.

Expand Down
44 changes: 44 additions & 0 deletions docs/compliance.md
Original file line number Diff line number Diff line change
Expand Up @@ -127,6 +127,50 @@ Each `FrameworkReport` exposes `framework_key`, `framework_name`, `passed` (bool
`audit_entries` or the HIPAA identifier coverage), plus its own `to_dict()`,
`to_json()`, and `to_frame()`.

## Pseudonymisation keys {#pseudonymisation-keys}

The privacy policy engine (`freshdata.enterprise.apply_privacy_policy`) ships
HIPAA, FERPA, PCI and GDPR packs. Their `pseudonymize` rules (the GDPR pack's
default action, the HIPAA date-of-birth rule and the FERPA grade rule) are
keyed, as are the `tokenize`, `surrogate` and `fpe` strategies of
`MaskingRule` in `anonymize` and `clean_enterprise`.

Pass a secret key for stable, joinable output, preferably from the environment:

```python
from freshdata.enterprise import PrivacyPolicy, apply_privacy_policy, load_compliance_pack

policy = PrivacyPolicy(
packs=(load_compliance_pack("gdpr"),),
jurisdiction="EU",
key_env="FRESHDATA_PSEUDONYM_KEY",
)
out, report = apply_privacy_policy(df, policy)
```

Without a key, each call uses a random key and emits `EphemeralKeyWarning`
(importable from `freshdata.enterprise`), and
`report.metadata["ephemeral_key_rules"]` lists the rules that used it. The
output cannot be recomputed from FreshData's source, but it changes on every
call, so it cannot be joined across runs. Policy `tokenize` without a key still
raises `ValueError`, as do `reversible=True` `tokenize` / `fpe` masking rules.

### Reproducing output from before 2.1.0

Before 2.1.0 these paths used constants from the public source when no key was
set. Anyone with that output and a list of candidate values can recompute the
pseudonyms, so treat it as reversible and re-pseudonymise it with a secret key.
If you must reproduce the old output for a while (for example, to join against
an existing table during a migration), pass the old constant as an explicit
key. **These keys are public: never use them for new data.**

| Before 2.1.0 (no key) | Explicit key that reproduces it |
| --- | --- |
| `MaskingRule(name=rule_name, strategy="tokenize")` | `key=hmac.new(b"freshdata-default-token-salt", rule_name.encode(), hashlib.sha256).hexdigest()[:32]` |
| `MaskingRule(strategy="surrogate")` | `key="freshdata-surrogate"` |
| `MaskingRule(strategy="fpe")` | `key="freshdata-surrogate"`, only when `pyffx` is not installed (with `pyffx`, a keyed `fpe` rule uses real FPE) |
| policy `pseudonymize` | `PrivacyPolicy(..., key="freshdata-surrogate")`, only when `pyffx` is not installed (a keyed `pseudonymize` uses FPE when `pyffx` is available) |

## Errors {#errors}

- `ValueError` — an unknown framework key was requested.
Expand Down
16 changes: 14 additions & 2 deletions docs/learning-profiles.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,8 +80,20 @@ mode, schema, and provenance.
## Privacy

By default (`privacy="mask"`), columns detected as sensitive — email, phone,
person name, national ID, address, postal code, or free text — never carry
raw literals in the saved profile:
person name, national ID (including medical record, patient, insurance and
driver's licence numbers), address, postal code, payment card number, bank
account / IBAN, IP address, health code, date of birth, or free text — never
carry raw literals in the saved profile. A column is sensitive when its name
matches a hint (`card_number`, `iban`, `acct`, `dob`, `email`, …; short hints
such as `pan` and `acct` must be a whole `_`-separated word) or when the
enterprise PII scanner finds any PII type in its values. A type without a
specific mapping is treated as free text rather than ignored; a date only
counts as a date of birth when the column name says so.

`freshdata profile audit` re-scans stored literals and exits `1` when a
profile that claims no raw values holds checksum-valid card numbers or IBANs
(profiles learned before these types were masked). Re-learn such a profile
and delete the old file.

- Rule-level evidence (a phone region, a `dayfirst` flag, a sentinel list)
carries no literals to begin with and replays normally.
Expand Down
6 changes: 6 additions & 0 deletions docs/parsers.md
Original file line number Diff line number Diff line change
Expand Up @@ -28,6 +28,12 @@ domain-validation are a separate step (`fd.clean`), so parsing and rules stay de
| SDMX-ML | `sdmx` | `observations` | — (audit-only) |
| UN/EDIFACT | `edifact` | `segments` | — |

GPX and SDMX documents may be UTF-8, UTF-16 or UTF-32 (with or without a byte-order
mark) or any encoding named in the XML declaration that the standard library reads.
Input over 10 MB, EBCDIC documents, and documents with a `DOCTYPE` or entity
declaration in any encoding are refused with an `unsafe ... XML` warning and empty
frames.

### FHIR R4 JSON

`fd.parse_domain(source, format="fhir")` accepts a **Bundle**, a single resource, a list
Expand Down
2 changes: 1 addition & 1 deletion docs/production-readiness.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,7 @@ clean data nobody is watching. Each item links to the relevant guarantee on the

## Install & pin

- [ ] Pin an exact version (`freshdata-cleaner==2.0.0`) and the extras you use
- [ ] Pin an exact version (`freshdata-cleaner==2.1.0`) and the extras you use
(`freshdata-cleaner[polars,privacy]`). Cleaning defaults can tighten between
minor versions — pinning keeps decisions reproducible.
- [ ] Install only the extras you need. The base install has **no** heavy deps;
Expand Down
Loading
Loading