Skip to content

fix(stream): reject unknown --timestamp columns - #426

Merged
kevincostner17 merged 1 commit into
mainfrom
fix/stream-cli-missing-timestamp
Sep 15, 2026
Merged

kevincostner17 merged 1 commit into
mainfrom
fix/stream-cli-missing-timestamp

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

freshdata stream --timestamp <column> accepted a column name that wasn't in the input. It exited 0 with no warning, and time-series cleaning was silently skipped for every batch, so late-data handling, interpolation and anomaly detection never ran.

The stream CLI now checks its column options against the first batch. An unknown name exits 1 with a one-line error, a did-you-mean hint when a column name is close, and no output file:

freshdata: error: --timestamp column 'event_tim' not found in input (did you mean 'event_time'?); input columns: event_time, value

The same check covers the other column options:

  • On stream, when --timestamp is given: --watermark, --entity-id and --ordered-dedupe-keys.
  • On stream, always: --target-column and --id-columns.
  • On stream-kafka: --target-column and --id-columns.

Root cause

When the timestamp column is missing, the time-series step returns the batch unchanged. That is reasonable for a single odd batch, but nothing checked the option against the input first, so a typo turned time-series mode off for the whole run with no message. --target-column and --id-columns had the same gap on the stream commands.

The check lives in streaming/_cli.py:

  • It runs on the first batch, before anything is cleaned or written.
  • A bad name raises ValueError, which the freshdata entry point already reports as freshdata: error: … with exit 1.
  • Nothing has been written at that point, so no output or <output>.partial file is created.
  • stream-kafka now reads from kafka_batches and passes the batches to clean_batches, so it gets the same check on its first raw batch.

Default-output changes

  • freshdata stream now exits 1 instead of 0 when any of these names a column that doesn't exist:

    • --timestamp
    • --watermark, --entity-id or --ordered-dedupe-keys (only when --timestamp is given)
    • --target-column or --id-columns

    In that case it prints a one-line freshdata: error: message on stderr and writes no output file, no per-batch reports and no summary.json.

  • freshdata stream-kafka now exits 1 instead of 0 when --target-column or --id-columns names a column missing from the first batch.

  • Runs where every named column exists produce the same output as before.

  • Without --timestamp, the time-series column options are still not checked, because they aren't used.

Tests

New file tests/test_streaming_cli_columns.py:

  • Unknown --timestamp (CSV and Parquet output):
    • exit 1 with a single freshdata: error: line naming the option and column
    • no output, .partial, per-batch report or summary.json
  • Hints:
    • a close name gets a did-you-mean hint
    • with no close match there's no hint, only the input column list
  • Valid --timestamp, with --entity-id and --watermark, still streams, and the summary includes the time-series section.
  • Other options: an unknown --watermark, --entity-id, --ordered-dedupe-keys, --target-column or --id-columns each exits 1 with no output.
  • Multiple bad options are reported together on one line.
  • No --timestamp: the time-series options aren't checked.
  • stream-kafka, with a mock consumer:
    • unknown --id-columns exits 1 with a hint and no output
    • valid columns still work

Verification

  • ruff check .: all checks passed.
  • mypy src/freshdata: no issues in 205 source files.
  • pytest -m "not online and not large":
    • Python 3.12 / pandas 2.3.3: 6317 passed, 14 skipped.
    • Python 3.9 / pandas 1.5.3: 6286 passed, 18 skipped. Two batch-cleaner timing tests (test_clean_duration_within_baselines on movies_json) failed while another suite was running on the same machine. They passed when rerun alone and don't exercise the stream CLI.

freshdata stream accepted a --timestamp that named no input column: it
exited 0 with no warning and time-series cleaning was quietly skipped for
every batch. The stream CLI now checks its column-name flags (--timestamp,
and with it --watermark, --entity-id and --ordered-dedupe-keys, plus
--target-column and --id-columns) against the first batch before anything
is cleaned or written. An unknown name exits 1 with a one-line
"freshdata: error: ..." message, a did-you-mean hint when one is close,
and no output or .partial file. stream-kafka checks --target-column and
--id-columns the same way.
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: e044e452-5d13-46f9-95a2-c566adebedf1


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit bf19ba7 into main Sep 15, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant