Skip to content

fix(memory): explicit sqlite memory selection, NaN-safe decisions, merged-profile schema, strict profile manifests - #394

Merged
kevincostner17 merged 1 commit into
mainfrom
fix/memory-learning-persistence
Sep 15, 2026
Merged

kevincostner17 merged 1 commit into
mainfrom
fix/memory-learning-persistence

Conversation

@kevincostner17

Copy link
Copy Markdown
Contributor

Summary

load_cleaning_memory (#306): a .sqlite/.db store keeps one row per
dataset_id, but loading ran "SELECT payload ... LIMIT 1" and returned
whichever row came first. It now takes a keyword-only dataset_id. A store
holding several memories without dataset_id raises ValueError listing the
stored ids; an unknown dataset_id raises KeyError. The store is opened
read-only via a file: URI (mode=ro), and a missing path raises
FileNotFoundError before connecting, so a mistyped path no longer creates
an empty database.

Decisions with NaN columns (#309): a decisions table round-tripped through
CSV has column=NaN for table-level steps. Replay keyed that as "nan", so
steps such as drop_duplicates never replayed. _normalize_decisions now maps
NaN/NaT/pd.NA and blank-string columns to None. CleaningMemory.to_json
replaces non-finite floats with null before json.dumps, so it never emits
bare NaN/Infinity tokens.

Merged LearningProfile drift (#308): merge built the audit with
alignment={"merged": True} and dropped source_schema. Replay then fell back
to the embedded memory signature's coarse type names ("text", "integer")
and compared them with pandas dtypes, reporting mild drift on the training
frame and excluding text columns from replay. The merged alignment now
carries the union of the parents' source_schema, and the preferred parent's
dtype wins a disagreement. The memory-signature fallback in _profile_schema
maps coarse types to pandas dtype names (text->object, integer->int64,
float->float64, boolean->bool, datetime->datetime64[ns]).

Strict .fdprofile manifests (#311): a truncated or non-object manifest
raised raw JSONDecodeError/KeyError, and hash verification only covered the
members the manifest listed, so an emptied member_hashes skipped
verification. Manifest decode/parse failures (ValueError, KeyError,
TypeError) and malformed member payloads now raise ProfileFormatError with
the original error chained. member_hashes must cover every required
non-manifest member, plus examples_vectors.npz when present. The hash
algorithm is unchanged, and profiles written by save_profile, including
ones with vectors, still load.

Verification

  • ruff check .: passed
  • mypy src/freshdata: no issues (202 files)
  • pytest -m "not online and not large" on main 2d098cb + this commit: Python 3.12 / pandas 2.3.3: 5160 passed, 13 skipped; Python 3.9 / pandas 1.5.3: 5156 passed, 17 skipped
  • New tests/test_memory_learning_persistence.py carries each issue's reproduction.

Closes #306
Closes #308
Closes #309
Closes #311

…rged-profile schema, strict profile manifests

load_cleaning_memory (#306): a .sqlite/.db store keeps one row per
dataset_id, but loading ran "SELECT payload ... LIMIT 1" and returned
whichever row came first. It now takes a keyword-only dataset_id. A store
holding several memories without dataset_id raises ValueError listing the
stored ids; an unknown dataset_id raises KeyError. The store is opened
read-only via a file: URI (mode=ro), and a missing path raises
FileNotFoundError before connecting, so a mistyped path no longer creates
an empty database.

Decisions with NaN columns (#309): a decisions table round-tripped through
CSV has column=NaN for table-level steps. Replay keyed that as "nan", so
steps such as drop_duplicates never replayed. _normalize_decisions now maps
NaN/NaT/pd.NA and blank-string columns to None. CleaningMemory.to_json
replaces non-finite floats with null before json.dumps, so it never emits
bare NaN/Infinity tokens.

Merged LearningProfile drift (#308): merge built the audit with
alignment={"merged": True} and dropped source_schema. Replay then fell back
to the embedded memory signature's coarse type names ("text", "integer")
and compared them with pandas dtypes, reporting mild drift on the training
frame and excluding text columns from replay. The merged alignment now
carries the union of the parents' source_schema, and the preferred parent's
dtype wins a disagreement. The memory-signature fallback in _profile_schema
maps coarse types to pandas dtype names (text->object, integer->int64,
float->float64, boolean->bool, datetime->datetime64[ns]).

Strict .fdprofile manifests (#311): a truncated or non-object manifest
raised raw JSONDecodeError/KeyError, and hash verification only covered the
members the manifest listed, so an emptied member_hashes skipped
verification. Manifest decode/parse failures (ValueError, KeyError,
TypeError) and malformed member payloads now raise ProfileFormatError with
the original error chained. member_hashes must cover every required
non-manifest member, plus examples_vectors.npz when present. The hash
algorithm is unchanged, and profiles written by save_profile, including
ones with vectors, still load.

Closes #306
Closes #308
Closes #309
Closes #311
@coderabbitai

coderabbitai Bot commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

Important

  • 🔍 Trigger review

This repository does not receive automatic reviews because it has fewer than 10 stars.

⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: 7dbc1a90-e5b5-41ee-986d-c67fafa10bfb


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown

FreshData benchmark report — performance

  • freshdata: ?
  • python: ?
  • platform: ?
fixture n_rows n_cols p50 s p95 s peak MB repair % false-repair % preserve % trust monotonic export %

Authored-code reduction (Metric 6)

@kevincostner17
kevincostner17 merged commit 96ae2fd into main Sep 15, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment