Skip to content

fix(semantic): use the validator phone and email patterns in type inference - #424

Merged
kevincostner17 merged 1 commit into
mainfrom
fix/semantic-types-phone-email
Sep 15, 2026
Merged

kevincostner17 merged 1 commit into
mainfrom
fix/semantic-types-phone-email

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

Semantic type inference (fd.infer_roles → semantic_type) now detects emails and phones with the same patterns as fd.validate_fields. This is a follow-up to #318.

Root cause

#318 fixed the email and phone validators in fieldcheck.py: punycode TLDs (xn--…), longer formatted international numbers, and a 7–15 digit count per E.164. semantic/semantic_types.py kept its own copies of the old regexes. As a result, inference returned unknown for user@example.xn--p1ai and +49 (0) 30 - 1234 - 0001. It also had no 15-digit cap, so 16–17 digit numbers could be inferred as phone.

Behaviour change

semantic_types now imports _EMAIL_RE and _is_phone from fieldcheck, so both accept exactly the same values. There is no import cycle, and fieldcheck is unchanged.

  • Emails with punycode TLDs are detected as email.
  • Formatted international phone numbers longer than 17 characters, with 7–15 digits, are detected as phone.
  • Numbers with more than 15 digits are no longer phone.
  • These stay unchanged:
    • the other detectors
    • the role="id"/role="text" short-circuits
    • the distinct-support gate
    • the name-hint veto
  • Numeric strings with 7–15 digits are still phone, as before and as in validate_fields.

Default-output changes

Only the semantic_type, semantic_type_confidence and semantic_type_evidence columns from fd.infer_roles can change, along with the same columns in ExplainReport.roles from explain_clean. Columns whose role is id or text are resolved before any content detector, so they are unaffected. For columns whose role is categorical or numeric:

  • Punycode-TLD email columns: category_code (0.4) → email (0.9).
  • Long formatted international phone columns, e.g. +44 (0) 20 7946 0958 or +49-(0)-30-1234-0001: category_code (0.4) → phone (0.9).
  • Columns of 16–17 digit numbers:
    • string columns: phone (0.9) → category_code (0.4)
    • int64 columns: phone (0.9) → unknown (0.2)

No cleaning, semantic repair or masking changes. The default cleaning path does not use semantic_type, and learn_memory and the compliance adapter read only role, missing_pct and domain_sensitive from infer_roles.

Tests

New tests in tests/test_semantic_types.py fail on main and pass with this change:

  • both repros, with 40 distinct values each
  • the 15-digit cap: 15 digits is phone; 16, 17 and 20 are not
  • ASCII emails and US phone formats are still detected
  • numeric ID columns are not phone: short IDs, 18-digit string and int64 IDs, and role="id"
  • infer_roles end to end on categorical columns
  • a parity test: for 26 sample values, the email and phone verdicts match whether validate_fields accepts the value

Verification

  • ruff check .: passed
  • mypy src/freshdata: no issues
  • pytest -m "not online and not large": Python 3.12 / pandas 2.3.3: 6289 passed, 14 skipped; Python 3.9 / pandas 1.5.3: 6260 passed, 18 skipped
  • pytest tests/truthbench: 290 passed

…erence

Semantic type inference kept its own email and phone regexes after #318
fixed the ones in fieldcheck. Punycode TLD emails and long international
phone formats were inferred as non-email/non-phone, and numbers with more
than 15 digits could be inferred as phone.

Reuse fieldcheck's _EMAIL_RE and _is_phone so inference accepts exactly
what validate_fields accepts. fieldcheck is unchanged.
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: a554662a-724d-4ecd-b74e-3a56698fae3f


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit bbbd93b into main Sep 15, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant