Skip to content

fix(semantic): suggest instead of auto-applying 24:00 -> 00:00 in time-only columns - #398

Merged
kevincostner17 merged 1 commit into
mainfrom
fix/semantic-time-24h-suggest
Sep 15, 2026
Merged

kevincostner17 merged 1 commit into
mainfrom
fix/semantic-time-24h-suggest

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

TimeCanonicalExpert proposed rewriting '24:00' to '00:00' at base confidence 0.96, above the 0.95 auto threshold, so semantic_mode='auto' and 'review' both applied it. The rationale, "the instant is unchanged", is only true when the value has a date: 24:00 on day D is 00:00 on day D+1. In a column of clock times alone, the rewrite turns end-of-day into start-of-day, so a shift_end of 24:00 sorted before its shift_start.

  • The proposal now scores 0.80, between the review (0.70) and auto (0.95) thresholds. It is recorded as a suggestion and routed for human review in every mode, and '24:00' is kept.
  • The rationale and docstring now explain why it is held for review.
  • Suggestions are not learned into cleaning memory. The time_canonical memory-replay case is replaced by a test that nothing is learned or replayed.
  • TruthBench oracle change (logistics):
    • log-06 transport_time '24:00' moves from REPAIR to REVIEW.
    • It was the domain's only REPAIR cell, and every domain must cover all dispositions, so log-07 tracking_status 'on time' → 'on-time' is added as the domain's exact repair. It is repaired automatically with full audit evidence.

Tests

  • New tests/test_semantic_time_24h.py:
    • Issue repro in auto, review and assist modes: '24:00' is kept, the status is suggested, and the rationale is the new text.
    • The proposal's confidence lies between the review and auto thresholds for both 24:00 and 24:00:00.
  • tests/test_semantic_repair_safety.py: a time_canonical suggestion is not learned and does not replay.
  • tests/truthbench/test_fixtures.py: log-06 is REVIEW; the new log-07 REPAIR cell and its expected output are checked; the family list is updated.

Verification

  • ruff check .: passed; mypy src/freshdata: no issues
  • pytest -m "not online and not large": Python 3.12: 5140 passed, 13 skipped; Python 3.9 / pandas 1.5: 5136 passed, 17 skipped
  • pytest tests/truthbench: 250 passed
  • python -m benchmarks.truthbench run --profile release --backends pandas,polars,duckdb --require-backends --repeats 2 --check: all 48 gates pass (also 48/48 on main 2d098cb)

Closes #305

…e-only columns

TimeCanonicalExpert proposed '24:00' -> '00:00' at base_confidence 0.96,
above the 0.95 auto threshold, so semantic_mode='auto' and 'review' both
applied it. Its rationale, "the instant is unchanged", only holds when a
date travels with the value: 24:00 on day D is 00:00 on day D+1. In a
column of clock times alone the rewrite turns end-of-day into
start-of-day, so a shift_end of 24:00 sorted before its shift_start.

The proposal now scores 0.80, between the review (0.70) and auto (0.95)
thresholds, so it is recorded as a suggestion and routed to a human in
every mode. The rationale and docstring say why it is held for review.

Suggestions are never learned into cleaning memory, so the replay case
for time_canonical is replaced by a test that nothing is learned or
replayed. The TruthBench logistics oracle for log-06 transport_time moves
from REPAIR to REVIEW. That was the domain's only REPAIR cell and every
domain must cover all dispositions, so log-07 tracking_status 'on time'
-> 'on-time' is added as the domain's exact repair.

Closes #305
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: a6616dac-9bbe-4347-a98b-c41340982b81


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit 5a73aad into main Sep 15, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

semantic_mode='auto' rewrites '24:00' to '00:00' in date-less time columns

1 participant