An experimental memory runtime for AI agents , built from scratch to understand what it actually means for a machine to remember.
Yes, I named it ChunkdUp purely because it sounded cool in my head. There is no deeper meaning. Don't look for one.
I'm building this because I don't want to treat "AI memory" as a magical API call.
When we say an agent "remembers" something, several different problems are hiding underneath:
- What information is actually worth remembering?
- What should happen when the same information appears again?
- What happens when a new fact contradicts an old one?
- How do we persist memory without turning the database into a dump of everything the user ever said?
- How do we retrieve the right memory when the user's query is vague, incomplete, or indirect?
- How do we know whether the memory system is actually working?
ChunkdUp is my attempt to answer these questions by building the underlying pieces myself, testing them, and keeping track of where the system still fails.
V1 focuses primarily on memory formation, state management, persistence, and establishing a measurable retrieval baseline.
The core idea is simple:
Conversation
│
▼
Extract useful information
│
▼
Memory Decision
│
├── STORE
├── MERGE
├── UPDATE
└── IGNORE
│
▼
Persistent Memory
│
▼
Retrieval
│
▼
Relevant Context
Rather than treating every conversation as something that should automatically become memory, ChunkdUp makes the memory lifecycle explicit.
"I use Python."
↓
STORE
"I use Python."
↓
MERGE
"I switched my backend to Go."
↓
UPDATE
"It's raining today."
↓
IGNORE
The result is not just a collection of vectors.
It is a piece of state that can evolve over time.
ChunkdUp started as a collection of small experiments. Each lab explored a problem that eventually became part of the larger system.
How do we decide what information survives when context is limited?
Explored:
- context constraints
- keyword-based selection
- semantic similarity
- greedy budget packing
The first lesson was simple:
Matching words isn't the same as matching meaning.
How do we safely consume output from a probabilistic model?
Explored:
- structured JSON output
- schema validation
- regex fallback parsing
- intentional rejection
- handling malformed model output
The goal was not to make the LLM magically reliable.
It was to make the system safe when the LLM isn't reliable.
How should memory evolve when new information arrives?
This became the core of ChunkdUp's memory lifecycle.
The system explicitly handles:
STORE
MERGE
UPDATE
IGNORE
and maintains memory state rather than blindly accumulating new records.
This is also where the project moved from being a collection of experiments toward becoming an actual memory runtime.
How do we turn stored memory into useful context for an AI system?
This connected the pieces:
User Query
↓
Memory Retrieval
↓
Context Assembly
↓
LLM
↓
Structured Output
The system could now retrieve dynamic user state and use it as part of an AI interaction.
The current V1 architecture looks like this:
flowchart TD
UserQuery[User Query / Conversation]
subgraph MemoryFormation[Memory Formation]
Extractor[Memory Extractor]
Decision[Decision Engine]
Repo[(Memory Repository)]
Extractor -->|Extract facts| Decision
Decision -->|STORE / MERGE / UPDATE / IGNORE| Repo
end
UserQuery --> Extractor
UserQuery --> Retriever
subgraph Retrieval[Memory Retrieval]
Retriever[Hybrid Retriever]
Semantic[Semantic Retrieval]
Keyword[Keyword Retrieval]
Ranking[Ranking / Reranking]
Retriever --> Semantic
Retriever --> Keyword
Semantic --> Ranking
Keyword --> Ranking
end
Repo -->|Active Memories| Retriever
Ranking --> PromptBuilder[Context Assembly]
subgraph Generation[LLM Interaction]
PromptBuilder --> LLM[LLM]
LLM --> Parser[Output Parser]
Parser --> Validator[Output Validator]
Validator --> FinalOutput([Structured Output])
end
The important architectural separation is:
Memory Formation
≠
Memory Retrieval
≠
LLM Generation
Each stage has a different responsibility.
One of the main goals of V1 was to make the memory lifecycle testable rather than anecdotal.
Current policy benchmark:
54 / 54 scenarios passed
100% policy accuracy
The benchmark covers decisions such as:
- new information → STORE
- repeated information → MERGE
- changed information → UPDATE
- irrelevant information → IGNORE
The current retrieval benchmark establishes a baseline for V2:
MRR 0.7857
Mean Recall@K 0.8571
Mean Precision@K 0.3821
These numbers are intentionally treated as a baseline, not as a claim that retrieval is solved.
For example, the system can successfully retrieve relevant memories for queries such as:
"What programming languages do I use?"
"What editor or IDE do I prefer?"
"What is my tech stack?"
while still struggling with cases such as:
"What is the weather forecast today?"
where the correct behavior should be:
[]
rather than returning an unrelated memory.
That failure is important.
It tells us that retrieval is not simply:
"Find the most similar memory."
The system also needs to determine:
"Is any stored memory relevant enough to return at all?"
The biggest lesson from V1 is that AI memory is not one problem.
It is a chain of problems:
Should I remember this?
↓
What exactly should I remember?
↓
How does this memory change?
↓
How should I persist it?
↓
Which memories are relevant later?
↓
Is the retrieved memory actually relevant?
↓
How should the agent use it?
Getting one layer right does not automatically make the entire system reliable.
V1 established the memory lifecycle and gave us a working retrieval baseline.
The next stage is intentionally focused on semantic retrieval.
This is where the problem gets considerably harder.
Humans rarely write perfectly explicit queries.
We say things like:
"What was that database again?"
"What did we decide?"
"Use the same thing as last time."
"Where do I run this?"
"Okay, let's continue."
The relevant memory may not share obvious keywords with the query.
The query may also be:
- short
- vague
- incomplete
- indirect
- context-dependent
- ambiguous
- completely unrelated to stored memory
So V2 will investigate how ChunkdUp can distinguish between:
Relevant memory
vs.
Semantically similar but irrelevant memory
vs.
No relevant memory
The V2 research pipeline will explore these questions incrementally:
Query
↓
Query understanding
↓
Candidate generation
↓
Semantic retrieval
↓
Keyword / entity signals
↓
Ranking
↓
Relevance decision
↓
Final memories
Potential areas of experimentation include:
- contextual memory representations
- query rewriting
- entity-aware retrieval
- hybrid retrieval strategies
- adaptive retrieval depth
- relevance thresholds
- reranking
- out-of-domain rejection
- ambiguous and indirect queries
- evaluation across different query categories
The important rule for V2:
No new retrieval component gets added simply because it sounds interesting. It needs a concrete failure case and a benchmark that can tell us whether it actually helped.
ChunkdUp is intentionally being built incrementally.
I don't want the project to become a collection of every technique that exists in the AI-memory ecosystem.
The process is:
Build
↓
Measure
↓
Find a failure
↓
Understand the failure
↓
Change one thing
↓
Measure again
If something works, keep it.
If something fails, understand why.
If a more complicated architecture does not solve a demonstrated problem, don't add it.
- Memory extraction
- Deterministic memory policy
- STORE / MERGE / UPDATE / IGNORE
- Memory repository abstraction
- Persistent storage
- PostgreSQL + pgvector
- Memory history
- Optimistic locking
- Hybrid retrieval baseline
- Entity-aware retrieval experiments
- Evaluation harness
- Memory policy benchmark
- Retrieval benchmark
- Basic observability metrics
- Improve vague-query retrieval
- Improve contextual retrieval
- Improve out-of-domain rejection
- Investigate contextual memory representations
- Evaluate query rewriting
- Evaluate entity-aware retrieval
- Improve relevance filtering
- Benchmark retrieval failure categories
- Compare retrieval strategies systematically
I'm not trying to claim that ChunkdUp has solved AI memory.
Quite the opposite.
The interesting part is discovering where the simple solution stops working.
V1 gave me a working memory lifecycle and a measurable baseline.
V2 is about understanding the harder question:
When a human asks an AI something vague, how does the system figure out which piece of its past actually matters?
That's the next problem I'm going after.