fix(cli): JSON-only dbt-gate stdout, validated --config on every engine, errors to stderr - #428
Merged
kevincostner17 merged 5 commits intoSep 15, 2026
Conversation
evaluate_trust_gate called fd.clean with clean_config=None, so CleanConfig's verbose=True default printed the clean summary and its warnings to stdout ahead of dbt-gate's JSON summary, and 'dbt-gate | jq' failed. The shared gate (dbt, Airflow, Dagster) now cleans quietly when no config is given and logs the clean warnings through the freshdata.integrations logger; an explicit CleanConfig is still used as given. dbt-gate also routes anything printed while gating to stderr, so stdout carries only the summary.
_build_enterprise read only six scalar keys plus masking, semantic and clustering from the 'enterprise' section, and the loader never looked at top-level section names. Typos such as enable_maskin or fail_under_trus, a misspelled 'enterprize:' section, and real EnterpriseConfig fields such as enable_privacy_detection or privacy were all silently ignored with exit 0. Unknown sections and keys, including keys inside nested objects, now fail with a one-line did-you-mean error (exit 1), worded like the 'clean' section's error. Every EnterpriseConfig field is either applied or rejected: enable_privacy_detection, enable_entity_resolution, trust_weights, lineage, privacy, k_anonymity and entity_resolution are now built into the config; enable_contracts, drift and anonymization are rejected with the reason, because the clean command never runs contract checks or applies them. Boolean, actor and fail_under_trust values are type-checked. The accepted keys are documented in docs/feature-overview.md.
…ields
enable_contracts, drift and anonymization do nothing in freshdata clean, so
the previous commit rejected them outright. That also broke configs that only
spell out their defaults. A null or default value (enable_contracts: false,
drift: {} or an object of DriftConfig defaults, anonymization: []) is now
accepted and ignored. Any other value still exits 1 with the reason, and
unknown keys inside a drift object still get a did-you-mean error.
With --engine polars, duckdb, spark, freshcore or auto, cmd_clean returned into the engine path before reading --config, so the whole file was silently ignored: clean options did not apply, typos did not error, and enterprise options such as masking were dropped without a word. The config file is now loaded and validated before engine dispatch, so a typo errors on every engine. The engine path merges the 'clean' section under the command-line options, as the pandas path does, and rejects 'context'/'policy' there with a one-line error. Native engines do not run the enterprise stage, so an 'enterprise' section that sets anything other than the EnterpriseConfig defaults exits 1 and names the keys.
freshdata clean, profile, learn, plan, apply-plan, policy compile and models pull printed their own error and usage messages to stdout, so a pipeline reading stdout (for example --output-format json | jq) got an error line in its data. Examples: 'error: cannot load profile ...', 'error: Unknown model id ...'. Exit-1 errors and validate's exit-2 errors already went to stderr. These messages now go to stderr in the one-line 'freshdata: error: ...' form that main's top-level handler uses. Exit codes are unchanged. Tests that read these messages from stdout now read stderr.
Contributor
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
FreshData benchmark report —
|
| fixture | n_rows | n_cols | p50 s | p95 s | peak MB | repair % | false-repair % | preserve % | trust | monotonic | export % |
|---|
Authored-code reduction (Metric 6)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
This PR fixes four CLI output bugs.
dbt-gatestdout was not valid JSON. Clean summary and warning lines were printed before the JSON summary.freshdata clean --configsilently ignored parts of the file. Misspelledenterprisekeys, misspelled top-level sections, and realEnterpriseConfigfields the loader never read were all dropped with exit 0.freshdata clean --engine polars|duckdb|spark|freshcore|autoignored--configentirely. Its clean options were not applied, typos did not error, and enterprise options such as masking were dropped with no message.clean,profile,learn,plan,apply-plan,policy compileandmodels pullwere printed there, so an error line could end up in piped data.Root cause
dbt-gate:evaluate_trust_gate(integrations/_core.py) calledfd.clean(df, config=None, report=True).CleanConfig.verbosedefaults toTrue, sofd.cleanprintedfreshdata: rows …andwarning: …to stdout beforedbt-gatewrote its JSON.--configkeys:_build_enterprise(enterprise/cli.py) read nine keys, ignored every other key, and never checked top-level section names.--configon native engines:cmd_cleanhanded off to_cmd_clean_enginebefore the config file was read.print(). Onlymain's top-level handler (exit 1) andvalidatewrote to stderr.Behaviour change
Trust gate (dbt, Airflow, Dagster)
clean_config, the gate cleans quietly and logs clean warnings at WARNING on thefreshdata.integrationslogger.CleanConfigis used as given.dbt-gatealso redirects anything printed while gating to stderr, so stdout is exactly one JSON document.--configvalidationenterprisekeysenable_*must be a boolean,actora string or null, andfail_under_trusta number or null.Newly applied enterprise fields
enable_privacy_detection,enable_entity_resolution,trust_weights,lineage,privacy,k_anonymityandentity_resolutionare now built intoEnterpriseConfig.Fields that do nothing in
freshdata cleanenable_contracts,driftandanonymizationhave no effect inclean. They are accepted and ignored when null or at their default:enable_contracts: falsedrift: {}, or an object holding onlyDriftConfigdefaultsanonymization: []Any other value exits 1 with a reason:
cleantakes no baseline or contract, and no pipeline appliesanonymization.Native engines
--configis loaded and validated before engine dispatch, so a typo errors on every engine.cleansection is merged under the command-line options, as on pandas.contextorpolicyin thecleansection exits 1, because only pandas supports them.enterprisesection that sets anything other than the defaults exits 1 and names the keys, because native engines do not run the enterprise stage.Errors to stderr
enterprise/cli.pycommands now go to stderr as one line:freshdata: error: …. This matches the top-level handler.Docs
docs/feature-overview.mddocuments the accepted keys and the engine behaviour.docs/integrations.mdnotes thatdbt-gatestdout is JSON only.Default-output changes
dbt-gatestdoutfreshdata: rows A->B, cols …, missing …line and thewarning: …lines no longer appear there.freshdata: <warning>log records. With no logging configured, Python's last-resort handler prints them.Library trust gate with no
clean_configThis covers
evaluate_trust_gate, the Airflow operator, Dagster assets and resources,FreshDataDbtTransform.runandgate_manifest.freshdata.integrationsinstead.CleanConfigis unchanged.Errors moved from stdout to stderr (exit codes unchanged)
Each of these now prints a
freshdata: error: …line on stderr:Profile loading:
freshdata clean --profileandfreshdata profile audit|diff|mergefailures, ascannot load profile …orcannot read profile …. Previously these printederror: …on stdout.Pandas-only options on native engines:
freshdata clean --engine <native>with--context-fileor--profile.Strict policy: failures in
freshdata cleanandfreshdata policy compile.freshdata plan: policy errors and "no repair plan".freshdata apply-plan: plan-drift errors (exit 2) and protected-column errors (exit 3).freshdata learn: errors.freshdata profile merge: errors.freshdata models pull: errors.Usage messages:
profile audit,profile diffandprofile mergeusageprofile mergewithout-ofreshdata profile <data file>The
usage:lines had noerror:prefix before.New exit-1 errors in
freshdata clean --config(pandas engine)These files used to load with exit 0:
enterprisekey, e.g.enterprize:,enable_maskin,fail_under_trusenable_contracts,driftoranonymizationset to a non-default value (null and default values still load)privacylineagetrust_weightsk_anonymityentity_resolutiondriftenable_*value that is not a boolean, e.g."false", which used to count as trueactorthat is not a stringNew exit-1 errors in
freshdata clean --engine polars|duckdb|spark|freshcore|auto --configThese files used to be ignored with exit 0:
cleanorenterprisesections or in the top-level section namesenterprisesection that sets anything other than the defaults, e.g.masking,fail_under_trustorenable_clustering: truecontextorpolicyin thecleansectionChanged error text (still exit 1)
masking,semanticorclusteringentries now reportunknown 'masking[0]' key(s): 'colums' (did you mean 'columns'?). They previously reportedMaskingRule.__init__() got an unexpected keyword argument.fail_under_trustnow reports'fail_under_trust' must be a number or null.Config keys that now take effect (previously ignored)
Pandas engine:
enable_privacy_detectionwithprivacy: changes the output data. Detected PII is anonymized, and the summary gains aprivacy:line.k_anonymity: adds ak-anonymity (k=…)summary line.enable_entity_resolutionwithentity_resolution: adds anentity resolution (…)summary line.trust_weights: changes the reported trust scores. That can change thefail_under_trustgate result and the exit code.lineage: changes the lineage metadata written by--lineage.Native engines:
cleansection now applies. For example,column_names: falsekeeps the original column names.--strategyand--drop-duplicatesstill take precedence.Tests
tests/test_integrations/test_dbt.py:dbt-gatestdout parses withjson.loads.tests/test_integrations/test_core.py:CleanConfig(verbose=True)still prints.tests/test_cli_malformed_inputs.py:enterprisekeys, a misspelled top-level section, and unknown nested keysenable_contracts,driftandanonymization; adrifttypo still errorsEnterpriseConfigfield is handledenable_privacy_detection/privacy,lineage,k_anonymityandentity_resolutiontake effectclusteringobject still means no clusteringtests/test_execution/test_cli_engine.py:clean.contextis rejected.cleansection applies on polars and duckdb, and on spark when pyspark is installed.enterprisesection is accepted.tests/test_cli_error_streams.py:clean --profile, strict context, native-engine restrictions,profileusage/audit/diff/merge/extra arguments,policy compile --strictandmodels pullfreshdata: error:message on stderrtests/learning/test_cli.py,tests/test_cli_models.pyandtests/context/test_context_cli.py.--configtests pass unchanged.Verification
ruff check .: all checks passed.mypy src/freshdata: no issues in 205 source files.pytest -m "not online and not large":main, the CLI, integration and context test files were rerun: 355 passed, 1 skipped.dbt-gate --manifest … --conn sqlite:///wh.sqlite --threshold 0 | python -c "import json,sys; json.load(sys.stdin)"succeeds.freshdata clean in.csv -o out.csv --config typo.jsonexits 1 withunknown key(s): 'enable_maskin' (did you mean 'enable_masking'?), 'fail_under_trus' (did you mean 'fail_under_trust'?).freshdata clean in.parquet -o out.parquet --engine polars --config cfg.json:{"clean": {"column_names": false}}, keepsCustomer IDenterprise.maskingrule, exits 1 and namesmaskingfreshdata models pull no-such-modelexits 2 with empty stdout andfreshdata: error: …on stderr.Not in this PR
Native engines still ignore several
cleancommand-line flags without a message:--mask,--cluster,--fail-under-trust,--lineageand the semantic flags.--maskmatters most, because it returns raw values when masking was requested. It needs its own change.