Commit 567a30e
authored
test: cover file ingestion and round-trip safety for clean_csv/clean_excel (#478)
Adds tests/test_ingestion_roundtrip.py, pinning what survives the reader
boundary in fd.clean_csv and fd.clean_excel — the dtypes a caller gets back
are decided by pandas before any cleaning step runs.
Covered:
* Excel serial dates (45000) are read as plain integers; only a
date-formatted cell becomes a timestamp, and the column name is no hint.
* duplicate headers in CSV and Excel (pandas "name.1" -> "name_1"), plus
labels that collide only after normalization ("Name"/"name "/"NAME"),
and the stability of those names across a save and reload.
* cp1252/latin-1 input: the bare UnicodeDecodeError with no encoding
option, accented text preserved when an encoding is passed, C1 control
characters from a latin-1 misread, and mojibake that survives cleaning
untouched while fd.lint_text_encoding flags it.
* raw -> clean -> save -> reload -> clean for CSV and Excel, asserting the
second clean is a no-op down to the written bytes, including quoting,
embedded newlines and empty cells.
* formula sanitizing: the returned frame is never rewritten, the written
file is, so cycle 1 -> 2 changes values once and is idempotent after;
sanitize_formulas=False round-trips byte-exactly.
* malformed and degenerate files: ragged long/short rows, header-only and
empty inputs.
Two assertions pin behaviour that loses information and say so inline:
preserve_leading_zeros is a no-op for clean_excel (pandas turns the text
cell "02134" into 2134, so a CSV -> Excel hop undoes what clean_csv kept),
and the leading-zero pre-scan stops at LEADING_ZERO_SCAN_ROWS.1 parent bd3d6fe commit 567a30e
1 file changed
Lines changed: 413 additions & 0 deletions
0 commit comments