Skip to content

Commit 567a30e

Browse files
test: cover file ingestion and round-trip safety for clean_csv/clean_excel (#478)
Adds tests/test_ingestion_roundtrip.py, pinning what survives the reader boundary in fd.clean_csv and fd.clean_excel — the dtypes a caller gets back are decided by pandas before any cleaning step runs. Covered: * Excel serial dates (45000) are read as plain integers; only a date-formatted cell becomes a timestamp, and the column name is no hint. * duplicate headers in CSV and Excel (pandas "name.1" -> "name_1"), plus labels that collide only after normalization ("Name"/"name "/"NAME"), and the stability of those names across a save and reload. * cp1252/latin-1 input: the bare UnicodeDecodeError with no encoding option, accented text preserved when an encoding is passed, C1 control characters from a latin-1 misread, and mojibake that survives cleaning untouched while fd.lint_text_encoding flags it. * raw -> clean -> save -> reload -> clean for CSV and Excel, asserting the second clean is a no-op down to the written bytes, including quoting, embedded newlines and empty cells. * formula sanitizing: the returned frame is never rewritten, the written file is, so cycle 1 -> 2 changes values once and is idempotent after; sanitize_formulas=False round-trips byte-exactly. * malformed and degenerate files: ragged long/short rows, header-only and empty inputs. Two assertions pin behaviour that loses information and say so inline: preserve_leading_zeros is a no-op for clean_excel (pandas turns the text cell "02134" into 2134, so a CSV -> Excel hop undoes what clean_csv kept), and the leading-zero pre-scan stops at LEADING_ZERO_SCAN_ROWS.
1 parent bd3d6fe commit 567a30e

1 file changed

Lines changed: 413 additions & 0 deletions

File tree

0 commit comments

Comments
 (0)