fix(stream): reject unknown --timestamp columns - #426
Merged
Merged
Conversation
freshdata stream accepted a --timestamp that named no input column: it exited 0 with no warning and time-series cleaning was quietly skipped for every batch. The stream CLI now checks its column-name flags (--timestamp, and with it --watermark, --entity-id and --ordered-dedupe-keys, plus --target-column and --id-columns) against the first batch before anything is cleaned or written. An unknown name exits 1 with a one-line "freshdata: error: ..." message, a did-you-mean hint when one is close, and no output or .partial file. stream-kafka checks --target-column and --id-columns the same way.
Contributor
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
FreshData benchmark report —
|
| fixture | n_rows | n_cols | p50 s | p95 s | peak MB | repair % | false-repair % | preserve % | trust | monotonic | export % |
|---|
Authored-code reduction (Metric 6)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
freshdata stream --timestamp <column>accepted a column name that wasn't in the input. It exited 0 with no warning, and time-series cleaning was silently skipped for every batch, so late-data handling, interpolation and anomaly detection never ran.The stream CLI now checks its column options against the first batch. An unknown name exits 1 with a one-line error, a did-you-mean hint when a column name is close, and no output file:
The same check covers the other column options:
stream, when--timestampis given:--watermark,--entity-idand--ordered-dedupe-keys.stream, always:--target-columnand--id-columns.stream-kafka:--target-columnand--id-columns.Root cause
When the timestamp column is missing, the time-series step returns the batch unchanged. That is reasonable for a single odd batch, but nothing checked the option against the input first, so a typo turned time-series mode off for the whole run with no message.
--target-columnand--id-columnshad the same gap on the stream commands.The check lives in
streaming/_cli.py:ValueError, which thefreshdataentry point already reports asfreshdata: error: …with exit 1.<output>.partialfile is created.stream-kafkanow reads fromkafka_batchesand passes the batches toclean_batches, so it gets the same check on its first raw batch.Default-output changes
freshdata streamnow exits 1 instead of 0 when any of these names a column that doesn't exist:--timestamp--watermark,--entity-idor--ordered-dedupe-keys(only when--timestampis given)--target-columnor--id-columnsIn that case it prints a one-line
freshdata: error:message on stderr and writes no output file, no per-batch reports and nosummary.json.freshdata stream-kafkanow exits 1 instead of 0 when--target-columnor--id-columnsnames a column missing from the first batch.Runs where every named column exists produce the same output as before.
Without
--timestamp, the time-series column options are still not checked, because they aren't used.Tests
New file
tests/test_streaming_cli_columns.py:--timestamp(CSV and Parquet output):freshdata: error:line naming the option and column.partial, per-batch report orsummary.json--timestamp, with--entity-idand--watermark, still streams, and the summary includes the time-series section.--watermark,--entity-id,--ordered-dedupe-keys,--target-columnor--id-columnseach exits 1 with no output.--timestamp: the time-series options aren't checked.stream-kafka, with a mock consumer:--id-columnsexits 1 with a hint and no outputVerification
ruff check .: all checks passed.mypy src/freshdata: no issues in 205 source files.pytest -m "not online and not large":test_clean_duration_within_baselinesonmovies_json) failed while another suite was running on the same machine. They passed when rerun alone and don't exercise the stream CLI.