Skip to content

fix(cli): JSON-only dbt-gate stdout, validated --config on every engine, errors to stderr - #428

Merged
kevincostner17 merged 5 commits into
mainfrom
fix/cli-dbt-gate-json-and-enterprise-config
Sep 15, 2026
Merged

kevincostner17 merged 5 commits into
mainfrom
fix/cli-dbt-gate-json-and-enterprise-config

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

This PR fixes four CLI output bugs.

  • dbt-gate stdout was not valid JSON. Clean summary and warning lines were printed before the JSON summary.
  • freshdata clean --config silently ignored parts of the file. Misspelled enterprise keys, misspelled top-level sections, and real EnterpriseConfig fields the loader never read were all dropped with exit 0.
  • freshdata clean --engine polars|duckdb|spark|freshcore|auto ignored --config entirely. Its clean options were not applied, typos did not error, and enterprise options such as masking were dropped with no message.
  • Several exit-2 and exit-3 errors went to stdout. Error and usage messages from clean, profile, learn, plan, apply-plan, policy compile and models pull were printed there, so an error line could end up in piped data.

Root cause

  • dbt-gate: evaluate_trust_gate (integrations/_core.py) called fd.clean(df, config=None, report=True). CleanConfig.verbose defaults to True, so fd.clean printed freshdata: rows … and warning: … to stdout before dbt-gate wrote its JSON.
  • --config keys: _build_enterprise (enterprise/cli.py) read nine keys, ignored every other key, and never checked top-level section names.
  • --config on native engines: cmd_clean handed off to _cmd_clean_engine before the config file was read.
  • Error stream: these command handlers returned their own errors with a bare print(). Only main's top-level handler (exit 1) and validate wrote to stderr.

Behaviour change

Trust gate (dbt, Airflow, Dagster)

  • With no clean_config, the gate cleans quietly and logs clean warnings at WARNING on the freshdata.integrations logger.
  • An explicit CleanConfig is used as given.
  • dbt-gate also redirects anything printed while gating to stderr, so stdout is exactly one JSON document.

--config validation

  • These now fail with exit 1 and a one-line did-you-mean error:
    • unknown top-level sections
    • unknown enterprise keys
    • unknown keys inside nested objects and list entries
  • Value types are now checked: enable_* must be a boolean, actor a string or null, and fail_under_trust a number or null.

Newly applied enterprise fields

enable_privacy_detection, enable_entity_resolution, trust_weights, lineage, privacy, k_anonymity and entity_resolution are now built into EnterpriseConfig.

Fields that do nothing in freshdata clean

enable_contracts, drift and anonymization have no effect in clean. They are accepted and ignored when null or at their default:

  • enable_contracts: false
  • drift: {}, or an object holding only DriftConfig defaults
  • anonymization: []

Any other value exits 1 with a reason: clean takes no baseline or contract, and no pipeline applies anonymization.

Native engines

  • --config is loaded and validated before engine dispatch, so a typo errors on every engine.
  • The clean section is merged under the command-line options, as on pandas.
  • context or policy in the clean section exits 1, because only pandas supports them.
  • An enterprise section that sets anything other than the defaults exits 1 and names the keys, because native engines do not run the enterprise stage.

Errors to stderr

  • Exit-2 and exit-3 error and usage messages from the enterprise/cli.py commands now go to stderr as one line: freshdata: error: …. This matches the top-level handler.
  • Exit codes are unchanged.

Docs

  • docs/feature-overview.md documents the accepted keys and the engine behaviour.
  • docs/integrations.md notes that dbt-gate stdout is JSON only.

Default-output changes

dbt-gate stdout

  • Stdout now holds only the JSON summary. The freshdata: rows A->B, cols …, missing … line and the warning: … lines no longer appear there.
  • Clean warnings go to stderr as freshdata: <warning> log records. With no logging configured, Python's last-resort handler prints them.
  • Exit codes are unchanged.

Library trust gate with no clean_config

This covers evaluate_trust_gate, the Airflow operator, Dagster assets and resources, FreshDataDbtTransform.run and gate_manifest.

  • The clean summary is no longer printed to stdout.
  • Clean warnings are logged at WARNING on freshdata.integrations instead.
  • Output with an explicit CleanConfig is unchanged.

Errors moved from stdout to stderr (exit codes unchanged)

Each of these now prints a freshdata: error: … line on stderr:

  • Profile loading: freshdata clean --profile and freshdata profile audit|diff|merge failures, as cannot load profile … or cannot read profile …. Previously these printed error: … on stdout.

  • Pandas-only options on native engines: freshdata clean --engine <native> with --context-file or --profile.

  • Strict policy: failures in freshdata clean and freshdata policy compile.

  • freshdata plan: policy errors and "no repair plan".

  • freshdata apply-plan: plan-drift errors (exit 2) and protected-column errors (exit 3).

  • freshdata learn: errors.

  • freshdata profile merge: errors.

  • freshdata models pull: errors.

  • Usage messages:

    • profile audit, profile diff and profile merge usage
    • profile merge without -o
    • extra arguments to freshdata profile <data file>

    The usage: lines had no error: prefix before.

New exit-1 errors in freshdata clean --config (pandas engine)

These files used to load with exit 0:

  • an unknown top-level section or enterprise key, e.g. enterprize:, enable_maskin, fail_under_trus
  • enable_contracts, drift or anonymization set to a non-default value (null and default values still load)
  • an unknown key inside a nested object that used to be ignored:
    • privacy
    • lineage
    • trust_weights
    • k_anonymity
    • entity_resolution
    • a non-default drift
  • a non-object value in any of those nested fields
  • an enable_* value that is not a boolean, e.g. "false", which used to count as true
  • an actor that is not a string

New exit-1 errors in freshdata clean --engine polars|duckdb|spark|freshcore|auto --config

These files used to be ignored with exit 0:

  • any typo in the clean or enterprise sections or in the top-level section names
  • an enterprise section that sets anything other than the defaults, e.g. masking, fail_under_trust or enable_clustering: true
  • context or policy in the clean section

Changed error text (still exit 1)

  • Unknown keys in masking, semantic or clustering entries now report unknown 'masking[0]' key(s): 'colums' (did you mean 'columns'?). They previously reported MaskingRule.__init__() got an unexpected keyword argument.
  • A non-numeric fail_under_trust now reports 'fail_under_trust' must be a number or null.

Config keys that now take effect (previously ignored)

Pandas engine:

  • enable_privacy_detection with privacy: changes the output data. Detected PII is anonymized, and the summary gains a privacy: line.
  • k_anonymity: adds a k-anonymity (k=…) summary line.
  • enable_entity_resolution with entity_resolution: adds an entity resolution (…) summary line.
  • trust_weights: changes the reported trust scores. That can change the fail_under_trust gate result and the exit code.
  • lineage: changes the lineage metadata written by --lineage.

Native engines:

  • Every option in the clean section now applies. For example, column_names: false keeps the original column names.
  • Command-line --strategy and --drop-duplicates still take precedence.

Tests

  • tests/test_integrations/test_dbt.py:
    • On a sqlite warehouse, dbt-gate stdout parses with json.loads.
    • Output printed while gating lands on stderr.
  • tests/test_integrations/test_core.py:
    • With no config, the gate writes nothing to stdout and logs every clean warning.
    • An explicit CleanConfig(verbose=True) still prints.
  • tests/test_cli_malformed_inputs.py:
    • did-you-mean errors for unknown enterprise keys, a misspelled top-level section, and unknown nested keys
    • non-default values rejected and null/default values accepted for enable_contracts, drift and anonymization; a drift typo still errors
    • wrongly typed values error
    • a guard that every EnterpriseConfig field is handled
    • enable_privacy_detection/privacy, lineage, k_anonymity and entity_resolution take effect
    • an empty clustering object still means no clustering
  • tests/test_execution/test_cli_engine.py:
    • On polars, duckdb, spark, freshcore and auto, config typos error, enterprise features are rejected, and clean.context is rejected.
    • The clean section applies on polars and duckdb, and on spark when pyspark is installed.
    • A default-valued enterprise section is accepted.
  • tests/test_cli_error_streams.py:
    • covers 12 paths: clean --profile, strict context, native-engine restrictions, profile usage/audit/diff/merge/extra arguments, policy compile --strict and models pull
    • each case asserts the exit code, an empty stdout, and a freshdata: error: message on stderr
  • Existing tests:
    • Stdout assertions for these errors now read stderr, in tests/learning/test_cli.py, tests/test_cli_models.py and tests/context/test_context_cli.py.
    • All other existing --config tests pass unchanged.

Verification

  • ruff check .: all checks passed.
  • mypy src/freshdata: no issues in 205 source files.
  • pytest -m "not online and not large":
    • Python 3.12, pandas 2.3.3: 6328 passed, 15 skipped
    • Python 3.9, pandas 1.5.3: 6299 passed, 19 skipped
  • After rebasing onto the current main, the CLI, integration and context test files were rerun: 355 passed, 1 skipped.
  • Manual checks:
    • dbt-gate --manifest … --conn sqlite:///wh.sqlite --threshold 0 | python -c "import json,sys; json.load(sys.stdin)" succeeds.
    • freshdata clean in.csv -o out.csv --config typo.json exits 1 with unknown key(s): 'enable_maskin' (did you mean 'enable_masking'?), 'fail_under_trus' (did you mean 'fail_under_trust'?).
    • freshdata clean in.parquet -o out.parquet --engine polars --config cfg.json:
      • with {"clean": {"column_names": false}}, keeps Customer ID
      • with an enterprise.masking rule, exits 1 and names masking
    • freshdata models pull no-such-model exits 2 with empty stdout and freshdata: error: … on stderr.

Not in this PR

Native engines still ignore several clean command-line flags without a message: --mask, --cluster, --fail-under-trust, --lineage and the semantic flags. --mask matters most, because it returns raw values when masking was requested. It needs its own change.

evaluate_trust_gate called fd.clean with clean_config=None, so CleanConfig's
verbose=True default printed the clean summary and its warnings to stdout
ahead of dbt-gate's JSON summary, and 'dbt-gate | jq' failed.

The shared gate (dbt, Airflow, Dagster) now cleans quietly when no config is
given and logs the clean warnings through the freshdata.integrations logger;
an explicit CleanConfig is still used as given. dbt-gate also routes anything
printed while gating to stderr, so stdout carries only the summary.
_build_enterprise read only six scalar keys plus masking, semantic and
clustering from the 'enterprise' section, and the loader never looked at
top-level section names. Typos such as enable_maskin or fail_under_trus, a
misspelled 'enterprize:' section, and real EnterpriseConfig fields such as
enable_privacy_detection or privacy were all silently ignored with exit 0.

Unknown sections and keys, including keys inside nested objects, now fail
with a one-line did-you-mean error (exit 1), worded like the 'clean' section's
error. Every EnterpriseConfig field is either applied or rejected:
enable_privacy_detection, enable_entity_resolution, trust_weights, lineage,
privacy, k_anonymity and entity_resolution are now built into the config;
enable_contracts, drift and anonymization are rejected with the reason,
because the clean command never runs contract checks or applies them.
Boolean, actor and fail_under_trust values are type-checked. The accepted
keys are documented in docs/feature-overview.md.
…ields

enable_contracts, drift and anonymization do nothing in freshdata clean, so
the previous commit rejected them outright. That also broke configs that only
spell out their defaults. A null or default value (enable_contracts: false,
drift: {} or an object of DriftConfig defaults, anonymization: []) is now
accepted and ignored. Any other value still exits 1 with the reason, and
unknown keys inside a drift object still get a did-you-mean error.
With --engine polars, duckdb, spark, freshcore or auto, cmd_clean returned
into the engine path before reading --config, so the whole file was
silently ignored: clean options did not apply, typos did not error, and
enterprise options such as masking were dropped without a word.

The config file is now loaded and validated before engine dispatch, so a
typo errors on every engine. The engine path merges the 'clean' section
under the command-line options, as the pandas path does, and rejects
'context'/'policy' there with a one-line error. Native engines do not run
the enterprise stage, so an 'enterprise' section that sets anything other
than the EnterpriseConfig defaults exits 1 and names the keys.
freshdata clean, profile, learn, plan, apply-plan, policy compile and
models pull printed their own error and usage messages to stdout, so a
pipeline reading stdout (for example --output-format json | jq) got an
error line in its data. Examples: 'error: cannot load profile ...',
'error: Unknown model id ...'. Exit-1 errors and validate's exit-2 errors
already went to stderr.

These messages now go to stderr in the one-line 'freshdata: error: ...'
form that main's top-level handler uses. Exit codes are unchanged. Tests
that read these messages from stdout now read stderr.
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: f660c6fa-6640-40bd-a378-96d46462f8c5


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit a80d0bd into main Sep 15, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant