fix(memory): explicit sqlite memory selection, NaN-safe decisions, merged-profile schema, strict profile manifests - #394
Merged
Conversation
…rged-profile schema, strict profile manifests load_cleaning_memory (#306): a .sqlite/.db store keeps one row per dataset_id, but loading ran "SELECT payload ... LIMIT 1" and returned whichever row came first. It now takes a keyword-only dataset_id. A store holding several memories without dataset_id raises ValueError listing the stored ids; an unknown dataset_id raises KeyError. The store is opened read-only via a file: URI (mode=ro), and a missing path raises FileNotFoundError before connecting, so a mistyped path no longer creates an empty database. Decisions with NaN columns (#309): a decisions table round-tripped through CSV has column=NaN for table-level steps. Replay keyed that as "nan", so steps such as drop_duplicates never replayed. _normalize_decisions now maps NaN/NaT/pd.NA and blank-string columns to None. CleaningMemory.to_json replaces non-finite floats with null before json.dumps, so it never emits bare NaN/Infinity tokens. Merged LearningProfile drift (#308): merge built the audit with alignment={"merged": True} and dropped source_schema. Replay then fell back to the embedded memory signature's coarse type names ("text", "integer") and compared them with pandas dtypes, reporting mild drift on the training frame and excluding text columns from replay. The merged alignment now carries the union of the parents' source_schema, and the preferred parent's dtype wins a disagreement. The memory-signature fallback in _profile_schema maps coarse types to pandas dtype names (text->object, integer->int64, float->float64, boolean->bool, datetime->datetime64[ns]). Strict .fdprofile manifests (#311): a truncated or non-object manifest raised raw JSONDecodeError/KeyError, and hash verification only covered the members the manifest listed, so an emptied member_hashes skipped verification. Manifest decode/parse failures (ValueError, KeyError, TypeError) and malformed member payloads now raise ProfileFormatError with the original error chained. member_hashes must cover every required non-manifest member, plus examples_vectors.npz when present. The hash algorithm is unchanged, and profiles written by save_profile, including ones with vectors, still load. Closes #306 Closes #308 Closes #309 Closes #311
Contributor
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Advanced Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
FreshData benchmark report —
|
| fixture | n_rows | n_cols | p50 s | p95 s | peak MB | repair % | false-repair % | preserve % | trust | monotonic | export % |
|---|
Authored-code reduction (Metric 6)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
load_cleaning_memory (#306): a .sqlite/.db store keeps one row per
dataset_id, but loading ran "SELECT payload ... LIMIT 1" and returned
whichever row came first. It now takes a keyword-only dataset_id. A store
holding several memories without dataset_id raises ValueError listing the
stored ids; an unknown dataset_id raises KeyError. The store is opened
read-only via a file: URI (mode=ro), and a missing path raises
FileNotFoundError before connecting, so a mistyped path no longer creates
an empty database.
Decisions with NaN columns (#309): a decisions table round-tripped through
CSV has column=NaN for table-level steps. Replay keyed that as "nan", so
steps such as drop_duplicates never replayed. _normalize_decisions now maps
NaN/NaT/pd.NA and blank-string columns to None. CleaningMemory.to_json
replaces non-finite floats with null before json.dumps, so it never emits
bare NaN/Infinity tokens.
Merged LearningProfile drift (#308): merge built the audit with
alignment={"merged": True} and dropped source_schema. Replay then fell back
to the embedded memory signature's coarse type names ("text", "integer")
and compared them with pandas dtypes, reporting mild drift on the training
frame and excluding text columns from replay. The merged alignment now
carries the union of the parents' source_schema, and the preferred parent's
dtype wins a disagreement. The memory-signature fallback in _profile_schema
maps coarse types to pandas dtype names (text->object, integer->int64,
float->float64, boolean->bool, datetime->datetime64[ns]).
Strict .fdprofile manifests (#311): a truncated or non-object manifest
raised raw JSONDecodeError/KeyError, and hash verification only covered the
members the manifest listed, so an emptied member_hashes skipped
verification. Manifest decode/parse failures (ValueError, KeyError,
TypeError) and malformed member payloads now raise ProfileFormatError with
the original error chained. member_hashes must cover every required
non-manifest member, plus examples_vectors.npz when present. The hash
algorithm is unchanged, and profiles written by save_profile, including
ones with vectors, still load.
Verification
ruff check .: passedmypy src/freshdata: no issues (202 files)pytest -m "not online and not large"on main 2d098cb + this commit: Python 3.12 / pandas 2.3.3: 5160 passed, 13 skipped; Python 3.9 / pandas 1.5.3: 5156 passed, 17 skippedtests/test_memory_learning_persistence.pycarries each issue's reproduction.Closes #306
Closes #308
Closes #309
Closes #311