Skip to content

Commit 78b7790

Browse files
fix(privacy): missing values stay missing, categorical k-anonymity, duplicate labels, fpe audit metadata, NER status (#399)
* fix(privacy): keep pd.NA and NaT cells missing when masking The rule path and detection path of anonymize(), detect_pii() and apply_privacy_policy() treated a cell as missing only when it was None or a float NaN. pd.NA (nullable string, Int64 and boolean columns) and NaT fell through that check, were stringified to "<NA>"/"NaT" and then redacted, tokenized or pseudonymised like real values. Missing cells came back as placeholders or tokens and cells_changed counted them. Add _is_missing_scalar(), which accepts None and any scalar that pd.isna() reports as missing, and use it for the four null checks. Missing cells are now passed through unchanged before masking; the masking functions themselves are unchanged, so non-null output is identical to before. Tests cover string, Int64, boolean, datetime64[ns] and object columns holding pd.NA across every rule strategy, the tokenize, pseudonymize, redact and quarantine policy actions, the detection path and the three reproductions from the issue. Closes #243 * fix(privacy): ignore empty category combinations in check_k_anonymity check_k_anonymity grouped rows with groupby(..., dropna=False).size() and the default observed=False. For categorical quasi-identifiers pandas then emits every combination of categories, including combinations with no rows. Those size-0 groups set smallest_class_size to 0, inflated n_equivalence_classes, turned ok to False and were listed in high_risk_groups, while the same data as object dtype passed. The same report feeds clean_enterprise with KAnonymityConfig. Group on the quasi-identifier columns with categorical columns cast to object, so only observed value combinations form groups, and drop any size-0 group. observed=True is not used because it mishandles dropna=False on pandas 1.5. Missing quasi-identifier values still form their own class. Tests cover the reproduction, a categorical quasi-identifier holding NaN, unused categories, mixed categorical and object quasi-identifiers, and clean_enterprise with KAnonymityConfig. Closes #244 * fix(privacy): reject duplicate column labels in detect_pii and anonymize detect_pii and the detection pass of anonymize loop over frame.columns and read frame[col].dtype. For a duplicated label frame[col] returns a DataFrame rather than a Series, so both functions failed with an unhelpful AttributeError. A masking rule that selected a duplicated label had the same problem, and _resolve_columns listed the label once per occurrence. Raise a ValueError naming the duplicated labels instead: - detect_pii raises up front when any label is duplicated. - anonymize raises when detection is enabled and any label is duplicated, or when a rule resolves to a duplicated label. The check runs before any rule is applied. A rule that targets a unique column still works on a frame that has duplicate labels elsewhere, as long as detection is not enabled. Tests cover both reproductions, a rule aimed at a duplicated label, and rules on a unique column next to duplicated ones with detection absent or disabled. Refs #265 * fix(privacy): report fpe reversibility and mode from what each cell used anonymize() computed MaskingEvent.reversible from the rule alone, so an fpe rule with reversible=True marked every event reversible even when pyffx was unavailable and the cell was masked with the one-way surrogate fallback. metadata["fpe_mode"] was also overwritten per cell, so a column mixing real FPE and the surrogate fallback reported only the mode of the last cell processed. _apply_rule_column now sets each event's reversible flag from the mode _mask_one returned for that cell: true only for tokenize, or for fpe when the cell used crypto_fpe, and only when the rule asked for it. It returns per-mode cell counts instead of the last mode. anonymize() aggregates those counts: a single mode is still reported as the plain fpe_mode string, and more than one mode is reported as fpe_mode="mixed" with fpe_modes={column: {mode: count}}. Masked values are unchanged. Tests use pyffx=None and a stub pyffx that cannot encrypt some values to cover the fallback, mixed modes within a column and across columns, reversible tokenize, and a single-mode report that matches the previous output. Refs #281 * fix(privacy): report whether detect_pii's NER pass actually ran detect_pii(config=PIIDetectionConfig(use_ner=True)) wrote metadata={"ner": True} from the config flag alone. When presidio_analyzer was importable but AnalyzerEngine() raised (for example a missing language model), _get_presidio_analyzer swallowed the exception without caching it. The NER pass then contributed nothing, the report still claimed NER ran, no warning was emitted, and the engine constructor was retried for every cell. _get_presidio_analyzer now records the first failure as "ExceptionType: message" in _PRESIDIO_ERROR and does not retry it. detect_pii resolves the analyzer once per call when use_ner is set, skips the per-cell NER pass when it is unavailable and emits one UserWarning naming the error. The metadata reports the real state: ner (now the active flag), ner_requested, ner_active, and ner_error when NER was requested but could not run. The PIIDetectionConfig docstring no longer says the pass is skipped silently. Tests stub presidio_analyzer through sys.modules: an engine that fails to start (one warning, correct metadata, constructor called once across cells and calls), a working engine returning no results (ner_active True), a missing package (import error recorded), and NER not requested. Closes #282
1 parent 5a73aad commit 78b7790

4 files changed

Lines changed: 587 additions & 22 deletions

File tree

‎src/freshdata/enterprise/config.py‎

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -294,7 +294,10 @@ class PIIDetectionConfig:
294294
The fallback detector (regex + context keywords) needs no extra
295295
dependencies. When ``use_ner`` is set and the optional
296296
``freshdata-cleaner[privacy]`` extra (Presidio) is installed, an NER pass is
297-
layered on top; otherwise it is skipped silently.
297+
layered on top. If Presidio is not installed or its analyzer fails to start,
298+
:func:`~freshdata.enterprise.detect_pii` emits one ``UserWarning``, runs the
299+
regex/context detector alone and records ``ner_active=False`` plus the
300+
``ner_error`` in the report metadata.
298301
"""
299302

300303
enabled: bool = True

‎src/freshdata/enterprise/privacy.py‎

Lines changed: 114 additions & 20 deletions
Original file line numberDiff line numberDiff line change
@@ -34,6 +34,7 @@
3434
import json
3535
import os
3636
import re
37+
import warnings
3738
from abc import ABC, abstractmethod
3839
from collections.abc import Callable
3940
from dataclasses import dataclass, field
@@ -55,6 +56,16 @@
5556
_PREVIEW_LEN = 24
5657

5758

59+
def _is_missing_scalar(value: Any) -> bool:
60+
"""True for ``None`` and any scalar missing marker (``NaN``, ``pd.NA``, ``NaT``).
61+
62+
Nullable ``string``/``Int64``/``boolean`` columns hold ``pd.NA`` and datetime
63+
columns hold ``NaT``; both must be passed through like ``None`` rather than
64+
stringified to ``"<NA>"``/``"NaT"`` and masked as if they were values.
65+
"""
66+
return value is None or (pd.api.types.is_scalar(value) and bool(pd.isna(value)))
67+
68+
5869
# =====================================================================
5970
# Entity patterns, context keywords, and HIPAA/GDPR maps
6071
# =====================================================================
@@ -449,28 +460,60 @@ def _ner_entities(
449460

450461

451462
_PRESIDIO_ANALYZER: Any = None
463+
#: ``"ExceptionType: message"`` from the first failed Presidio start-up. Cached
464+
#: so a missing package or language model is not retried for every cell.
465+
_PRESIDIO_ERROR: str | None = None
452466

453467

454-
def _get_presidio_analyzer() -> Any: # pragma: no cover - requires optional Presidio
455-
global _PRESIDIO_ANALYZER
456-
if _PRESIDIO_ANALYZER is None:
468+
def _get_presidio_analyzer() -> Any:
469+
"""Return the shared Presidio analyzer, or ``None`` when it cannot start.
470+
471+
The first failure is recorded in :data:`_PRESIDIO_ERROR` and not retried.
472+
"""
473+
global _PRESIDIO_ANALYZER, _PRESIDIO_ERROR
474+
if _PRESIDIO_ANALYZER is None and _PRESIDIO_ERROR is None:
457475
try:
458476
from presidio_analyzer import AnalyzerEngine
459477

460478
_PRESIDIO_ANALYZER = AnalyzerEngine()
461-
except Exception:
462-
_PRESIDIO_ANALYZER = None
479+
except Exception as exc:
480+
_PRESIDIO_ERROR = f"{type(exc).__name__}: {exc}"
463481
return _PRESIDIO_ANALYZER
464482

465483

484+
def _duplicated_labels(frame: pd.DataFrame) -> list[Any]:
485+
"""Column labels that occur more than once, in first-seen order."""
486+
return list(dict.fromkeys(frame.columns[frame.columns.duplicated()]))
487+
488+
466489
def detect_pii(df: Any, *, config: PIIDetectionConfig | None = None) -> PIIScanReport:
467490
"""Scan the text columns of *df* for PII; return a :class:`PIIScanReport`.
468491
469492
Read-only. Only object/string columns are scanned. Raw matched substrings
470493
are redacted in the report unless ``config.redact_samples=False``.
494+
495+
Raises :class:`ValueError` when *df* has duplicate column labels, because a
496+
duplicated label does not identify a single column to scan.
471497
"""
472498
cfg = config or PIIDetectionConfig()
473499
frame = to_pandas(df)
500+
duplicated = _duplicated_labels(frame)
501+
if duplicated:
502+
raise ValueError(
503+
f"detect_pii requires unique column labels; duplicated: {duplicated}"
504+
)
505+
ner_active = False
506+
ner_error: str | None = None
507+
if cfg.use_ner:
508+
ner_active = _get_presidio_analyzer() is not None
509+
if not ner_active:
510+
ner_error = _PRESIDIO_ERROR or "presidio analyzer unavailable"
511+
warnings.warn(
512+
f"detect_pii: use_ner=True but the Presidio NER pass is unavailable "
513+
f"({ner_error}); only the regex/context detector ran",
514+
UserWarning,
515+
stacklevel=2,
516+
)
474517
entities: list[PIIEntity] = []
475518
scanned: list[str] = []
476519
for col in frame.columns:
@@ -479,19 +522,26 @@ def detect_pii(df: Any, *, config: PIIDetectionConfig | None = None) -> PIIScanR
479522
continue
480523
scanned.append(str(col))
481524
for row, value in series.items():
482-
if value is None or (isinstance(value, float) and pd.isna(value)):
525+
if _is_missing_scalar(value):
483526
continue
484527
text = str(value)
485528
cell_entities = detect_in_text(text, column=str(col), config=cfg)
486-
if cfg.use_ner:
529+
if ner_active:
487530
cell_entities = _merge_ner(cell_entities, _ner_entities(text, str(col), cfg))
488531
for e in cell_entities:
489532
e.metadata["row"] = int(row) if isinstance(row, (int, float)) else row
490533
entities.extend(cell_entities)
534+
metadata: dict[str, Any] = {
535+
"ner": ner_active,
536+
"ner_requested": bool(cfg.use_ner),
537+
"ner_active": ner_active,
538+
}
539+
if ner_error is not None:
540+
metadata["ner_error"] = ner_error
491541
return PIIScanReport(
492542
entities=entities,
493543
columns_scanned=tuple(scanned),
494-
metadata={"ner": bool(cfg.use_ner)},
544+
metadata=metadata,
495545
)
496546

497547

@@ -955,25 +1005,53 @@ def anonymize(
9551005
"detection_config=PIIDetectionConfig() to say what to mask."
9561006
)
9571007
frame = to_pandas(df).copy()
1008+
duplicated = _duplicated_labels(frame)
1009+
if duplicated:
1010+
# A duplicated label selects several columns at once, so neither the
1011+
# detection pass nor a rule aimed at it can address a single column.
1012+
if detection_config is not None and detection_config.enabled:
1013+
raise ValueError(
1014+
"anonymize requires unique column labels for PII detection; "
1015+
f"duplicated: {duplicated}"
1016+
)
1017+
for rule in rules:
1018+
targeted = _resolve_columns(rule, duplicated)
1019+
if targeted:
1020+
raise ValueError(
1021+
f"anonymize requires unique column labels; rule {rule.name!r} "
1022+
f"targets duplicated: {targeted}"
1023+
)
9581024
events: list[MaskingEvent] = []
9591025
changed_cols: list[str] = []
9601026
cells_changed = 0
9611027
metadata: dict[str, Any] = {}
9621028

1029+
fpe_modes: dict[str, dict[str, int]] = {}
9631030
for rule in rules:
9641031
key = _resolve_key(rule)
9651032
vault = _vault_for(rule)
9661033
for column in _resolve_columns(rule, list(frame.columns)):
9671034
if column not in frame.columns:
9681035
continue
969-
n, fpe_mode = _apply_rule_column(
1036+
n, mode_counts = _apply_rule_column(
9701037
frame, column, rule, key, vault, events, audit_include_pii
9711038
)
9721039
if n:
9731040
cells_changed += n
9741041
changed_cols.append(str(column))
975-
if fpe_mode:
976-
metadata["fpe_mode"] = fpe_mode
1042+
if mode_counts:
1043+
per_column = fpe_modes.setdefault(str(column), {})
1044+
for mode, count in mode_counts.items():
1045+
per_column[mode] = per_column.get(mode, 0) + count
1046+
1047+
# One mode overall keeps the plain mode string; a mix of modes (per cell,
1048+
# column or rule) is reported as "mixed" with per-column cell counts.
1049+
modes_used = {mode for per_column in fpe_modes.values() for mode in per_column}
1050+
if len(modes_used) == 1:
1051+
metadata["fpe_mode"] = next(iter(modes_used))
1052+
elif modes_used:
1053+
metadata["fpe_mode"] = "mixed"
1054+
metadata["fpe_modes"] = fpe_modes
9771055

9781056
entities_found = 0
9791057
if detection_config is not None and detection_config.enabled:
@@ -1059,7 +1137,8 @@ def _apply_rule_column(
10591137
vault: TokenVault,
10601138
events: list[MaskingEvent],
10611139
include_pii: bool,
1062-
) -> tuple[int, str | None]:
1140+
) -> tuple[int, dict[str, int]]:
1141+
"""Mask one column in place; return ``(cells_changed, {fpe_mode: cell_count})``."""
10631142
if rule.strategy == "drop":
10641143
n = int(frame[column].notna().sum())
10651144
_record_event(
@@ -1068,39 +1147,44 @@ def _apply_rule_column(
10681147
source="column", original="", masked="<dropped>", include_pii=include_pii,
10691148
)
10701149
frame.drop(columns=[column], inplace=True)
1071-
return n, None
1150+
return n, {}
10721151

1073-
reversible = rule.strategy in ("tokenize", "fpe") and rule.reversible
10741152
format_preserving = rule.strategy in ("fpe", "surrogate") or rule.preserve_format
10751153
if rule.strategy in ("tokenize", "fpe") and rule.reversible and not key:
10761154
raise ValueError(
10771155
f"masking rule {rule.name!r}: reversible {rule.strategy} requires key= or key_env="
10781156
)
10791157

1080-
fpe_mode: str | None = None
1158+
mode_counts: dict[str, int] = {}
10811159
series = frame[column]
10821160
entity_type = _entity_for_rule(rule, str(column))
10831161
n_changed = 0
10841162
new_values: list[Any] = []
10851163
for row, value in series.items():
1086-
if value is None or (isinstance(value, float) and pd.isna(value)):
1164+
if _is_missing_scalar(value):
10871165
new_values.append(value)
10881166
continue
10891167
original = str(value)
10901168
masked, mode = _mask_one(original, rule, key, vault)
10911169
if mode:
1092-
fpe_mode = mode
1170+
mode_counts[mode] = mode_counts.get(mode, 0) + 1
10931171
new_values.append(masked)
10941172
if masked != original:
10951173
n_changed += 1
1174+
# Only a vault token or real FPE can be reversed; the surrogate
1175+
# fallback of ``fpe`` is one-way, whatever the rule asked for.
1176+
reversible = rule.reversible and (
1177+
rule.strategy == "tokenize"
1178+
or (rule.strategy == "fpe" and mode == "crypto_fpe")
1179+
)
10961180
_record_event(
10971181
events, column=str(column), row=row, entity_type=entity_type, rule=rule,
10981182
strategy=rule.strategy, reversible=reversible,
10991183
format_preserving=format_preserving, source="column",
11001184
original=original, masked=masked, include_pii=include_pii,
11011185
)
11021186
frame[column] = pd.Series(new_values, index=series.index)
1103-
return n_changed, fpe_mode
1187+
return n_changed, mode_counts
11041188

11051189

11061190
def _mask_one(
@@ -1152,7 +1236,7 @@ def _anonymize_detected(
11521236
touched = False
11531237
new_values: list[Any] = []
11541238
for row, value in series.items():
1155-
if value is None or (isinstance(value, float) and pd.isna(value)):
1239+
if _is_missing_scalar(value):
11561240
new_values.append(value)
11571241
continue
11581242
text = str(value)
@@ -1248,7 +1332,17 @@ def check_k_anonymity(
12481332
raise ValueError("check_k_anonymity requires at least one quasi-identifier")
12491333

12501334
n_rows = len(frame)
1251-
sizes = frame.groupby(list(quasi_identifiers), dropna=False).size()
1335+
# Categorical keys would make groupby emit every category combination,
1336+
# including empty ones. Group on the observed values instead; observed=True
1337+
# is avoided because it mishandles dropna=False on pandas 1.5.
1338+
keys = frame[list(quasi_identifiers)]
1339+
categorical = {
1340+
col: object for col, dtype in keys.dtypes.items() if isinstance(dtype, pd.CategoricalDtype)
1341+
}
1342+
if categorical:
1343+
keys = keys.astype(categorical)
1344+
sizes = keys.groupby(list(quasi_identifiers), dropna=False).size()
1345+
sizes = sizes[sizes > 0]
12521346
n_classes = int(len(sizes))
12531347
smallest = int(sizes.min()) if n_classes else 0
12541348
violating_mask = sizes < k

‎src/freshdata/enterprise/privacy_policy.py‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -48,6 +48,7 @@
4848
MaskingEvent,
4949
PrivacyReport,
5050
TokenVault,
51+
_is_missing_scalar,
5152
_luhn_ok,
5253
detect_in_text,
5354
detokenize_value,
@@ -688,7 +689,7 @@ def apply_privacy_policy(
688689
used_vault = tok_vault
689690
reversible = rule is None or rule.reversible
690691
for value in series:
691-
if value is None or (isinstance(value, float) and pd.isna(value)):
692+
if _is_missing_scalar(value):
692693
new_values.append(value)
693694
continue
694695
original = str(value)

0 commit comments

Comments
 (0)