Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 21 additions & 2 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -16,7 +16,26 @@ jobs:
- uses: actions/setup-python@v5
with:
python-version: ${{ matrix.python }}
- name: Install (all extras + dev)
run: pip install ".[all,dev]"
- name: Install (all extras + dev + encrypted)
run: pip install ".[all,dev,encrypted]"
- name: Run the suite
run: python -m pytest -q

# The client-wiring code is full of per-OS paths (Claude Desktop, VS Code, Zed live in three
# different places) — exercise it where those branches actually run. Core + MCP only: the heavy
# embedding extras are covered by the Linux job, and the suite degrades gracefully without them.
test-os:
runs-on: ${{ matrix.os }}
strategy:
fail-fast: false
matrix:
os: [windows-latest, macos-latest]
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- name: Install (core + MCP + dev)
run: pip install ".[mcp,dev]" numpy
- name: Run the suite
run: python -m pytest -q
72 changes: 69 additions & 3 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,12 +8,64 @@ breaking changes only land in a major. Format loosely follows

## [Unreleased]

## [1.0.0] — 2026-07-05
## [1.0.0] — 2026-07-06

The 1.0 contract: the API that exists today is the API, locked under semver. No new features gate
this release — it consolidates what shipped and hardened since 0.1.1 and freezes the surface.
The 1.0 contract: the public surface — the `midas.Memory` SDK, the guard, the coding `is_forbidden`
gate, the MCP tools, and the `midas` CLI — is now locked under semver. This release consolidates and
hardens everything since 0.1.1 (the retrieval core, the governance guard, the control-plane, the audit
chain, the client tooling, and the TypeScript port) and freezes the API.

### Added
- **The continuity control-plane** (`midas/continuity.py` + MCP tools): **`resume`** — everything an
agent needs to pick up a session in ONE call (pinned directives, forbidden rules, current state,
what changed, open loops, unresolved conflicts; token-budgeted, prompt-ready); **`memory_conflicts`**
— live beliefs that contradict each other with neither superseding the other (the multi-agent
shared-memory failure mode), NLI-scored when available with a same-slot heuristic fallback, ranked
candidates only; **open loops** — `remember_commitment` / `open_loops` / `close_loop` make promised
work survive across sessions with an auditable promise→resolution chain. The injected agent policy
now starts sessions with `resume` and keeps promises via open loops.
- **Per-kind TTL retention**: `Memory.forget_expired({"chat": 30})`, `MIDAS_MCP_TTL="chat=30,note=90"`,
and `maintain(ttl=…)` — age-based retention next to value-based forgetting; user-confirmed,
standing, and supersession-chain records never expire silently.
- **Tamper-evident audit chain** in `SQLiteStore`: every put/delete/clear appends a hash-chained
entry (hashes only — never memory content); `midas audit [--json]` shows and verifies the chain and
`midas doctor` checks its integrity. `audit=False` opts out for perf-sensitive paths.
- **`midas import --from claude-md | cursorrules | jsonl`**: file-based agent memory (CLAUDE.md,
.cursorrules, exported JSONL) becomes first-class Midas records — sectioned, deduped, idempotent,
and governance-safe by default (imported rules are `observation` provenance and cannot authorize
guarded actions unless you pass `--confirmed`).
- **4 more clients in `midas init`**: VS Code (user `mcp.json`, `servers` key), Gemini CLI, Cline,
and Zed (`context_servers`) — each written in its own config schema; status/uninstall understand
them too.
- **HTTP bearer-token auth**: `midas serve --http --token <secret>` / `MIDAS_MCP_TOKEN` gates every
request on the shared-URL transport (constant-time compare, 401 + `WWW-Authenticate` otherwise).
- **`midas doctor --json`** — the machine-readable diagnosis, same envelope as the wiring receipt.
- **TypeScript port: optional local semantic embeddings.** `LocalEmbedder` (ONNX via the optional
`@huggingface/transformers`, bge-small by default) behind `MIDAS_MCP_EMBEDDER=local`, with a clean
announced fallback to the byte-parity hashing embedder; the TS `Memory` API is now async
(`remember`/`capture`/`recall`/`buildContext`/`forgetMatching` return promises).
- **CI on Windows and macOS** (core + MCP): the client-wiring code is full of per-OS paths that only
Linux exercised.
- **The Agent Continuity Bench measures the new control-plane**: `resume_fidelity` (the one-call pack
contains live state / forbidden rules / open loops and never presents superseded values or closed
loops as live), `conflict_detection`, and `conflict_precision` — all → 1.00, wired into
`midas bench` with the same all-green verdict rule.
- **Inspector views** for the control-plane: Conflicts (ranked contradiction pairs with per-side
forget), Open loops (close with a recorded resolution), and Audit log (chain verification + a
hash-only tail).
- **TypeScript parity for the control-plane and audit chain**: `resume`, `memoryConflicts`
(heuristic tier), open loops, `forgetExpired`/TTL, and the same hash-chained `audit_log` — the hash
is canonicalised to integer microseconds so **either runtime verifies chains written by the other**
(validated bidirectionally).
- **`midas init --claude-hook`**: a Claude Code SessionEnd hook (`midas hook capture-session`) that
offers each finished session's user turns to `capture` — memory even when the agent never calls the
tool, with the same no-LLM keep/skip policy, fail-open by construction, removed by
`midas uninstall`.
- **`midas import --from mem0 | zep`**: migrate Mem0 exports (distilled memories → facts with
original ids/attribution preserved) and Zep exports (facts/edges → facts, messages → chat turns).
- **Encryption at rest (opt-in)**: `SQLiteStore(key=…)` / `MIDAS_MCP_KEY` with the new `[encrypted]`
extra (SQLCipher). The file on disk is ciphertext; opening without the key fails; if the key is set
but the extra is missing Midas **fails closed** instead of silently writing plaintext.
- **`midas init --json` / `midas status --json` — a machine-readable client wiring receipt** (#15).
One command now yields a compact, pasteable proof of what was wired: per client its config path,
detected/wired/changed state, backup path, and skip reason; plus the memory DB, scope mode
Expand All @@ -31,6 +83,20 @@ this release — it consolidates what shipped and hardened since 0.1.1 and freez
recall finding the signal among 1,500 noise records.

### Changed
- **`midas inspect` redesigned end to end.** A real light theme (not an inverted dark one) alongside the
existing dark brand theme, toggle + system-preference default + persisted choice; a fixed, validated
categorical color system so every `kind` and `provenance` value keeps the *same* color everywhere in
the app (Overview bars, Browse tags, Project governance) instead of one monochrome accent for
everything; a real 30-day activity chart (SVG line + area, hover crosshair/tooltip) and a recency
(short/medium/long) stacked bar on Overview; grouped, badge-annotated navigation (Memory / Coding /
Governance) with a command palette (**⌘K**) and a **/** search shortcut; native `confirm()`/`prompt()`
replaced with in-app modal/toast components; a responsive layout down to phone width with a horizontal
pill nav. New endpoints backing it: `api_timeseries` (daily capture counts), `api_meta` (version/db/
embedder/audit-chain identity), and `by_tier` added to `api_overview`. Fixed a real routing bug in the
process: the old build never listened for `hashchange`, so browser back/forward and direct links to a
view silently did nothing.
- The TypeScript `Memory` API is async (`remember`/`capture`/`recall`/`buildContext`/
`forgetMatching` return promises) to support the optional ONNX embedder.
- **A bare `Memory()` now auto-selects real semantic recall.** With the `[local]` extra installed,
`Memory()` upgrades from the offline hashing embedder to `LocalEmbedder` automatically (override
with `MIDAS_EMBEDDER=hashing|local`; default `auto`). Fixes the out-of-the-box recall weakness an
Expand Down
61 changes: 48 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,8 +68,9 @@ safely** and **resume from cleanly** — which is where similarity search alone
| *"How do I speed up the transactions list?"* | the **prior fix** resurfaces, so the agent doesn't re-diagnose it | — |

These properties are measured, not asserted — the **[agent-memory bench suite](docs/agent-memory-benches.md)**
scores action-safety, decision-adherence, repeated-mistake avoidance, and adversarial **memory-safety**
across scripted multi-session projects. The safety eval blocks **10 / 10 adversarial attacks** (ASR
scores action-safety, decision-adherence, repeated-mistake avoidance, **resume fidelity**, **conflict
detection/precision** (live contradictions between agents found without over-flagging), and adversarial
**memory-safety** across scripted multi-session projects. The safety eval blocks **10 / 10 adversarial attacks** (ASR
**0.00**) — including a planted confirmation next to a prohibition, a confirmation for a *different* action,
a provenance-laundering supersession, and a cross-namespace approval — with **no over-blocking** (benign-pass
**1.00**). Deterministic, $0, no LLM. **Reproduce every number with one command:**
Expand Down Expand Up @@ -120,17 +121,27 @@ wired to, under which scope/policy, and which clients were skipped (config paths
contents). Paste it into a bug report, or let another agent verify the setup without scraping prose.

`midas init` creates **one shared memory** (`~/.midas/memory.sqlite3`) and points the MCP clients it
detects — **Claude Code, Codex, Cursor, Claude Desktop, Windsurf** — at it. So all your agents read and
write the **same** memory, autonomously, with no per-client paths to keep in sync.
detects — **Claude Code, Codex, Cursor, Claude Desktop, Windsurf, VS Code, Gemini CLI, Cline, Zed** —
at it. So all your agents read and write the **same** memory, autonomously, with no per-client paths to
keep in sync.

Prefer a single endpoint over per-client launches? Run one server and give your clients an **MCP URL**:

```bash
midas serve --http # → http://127.0.0.1:7077/mcp (one server, one memory, every client shares it)
midas serve --http --token <secret> # require `Authorization: Bearer <secret>` on every request
```

Keep Midas current with **`midas update`**. See your memory anytime with **`midas inspect`**.

Already carrying agent memory in files? **`midas import --from claude-md CLAUDE.md`** (or
`--from cursorrules`, `--from jsonl`, `--from mem0`, `--from zep`) turns those rules and exports into
first-class, recallable, governable memories — tagged with where they came from, idempotent on re-run.

Want memory even when the agent never calls `capture`? **`midas init --claude-hook`** installs a
Claude Code SessionEnd hook that offers each session's user turns to memory — Midas's no-LLM policy
still decides what is actually kept.

<details>
<summary><b>Manual setup</b> — any client, or to customize (click to expand)</summary>

Expand All @@ -148,8 +159,12 @@ by default — no path needed. The universal block:
| **Claude Desktop** | Settings → Developer → Edit Config (`claude_desktop_config.json`) — paste, restart |
| **Codex CLI** | `codex mcp add midas -- midas-mcp` |
| **Windsurf** | `~/.codeium/windsurf/mcp_config.json` — paste the block |
| **VS Code** | user `mcp.json` (`servers` key, `"type": "stdio"`) — `midas init` writes it |
| **Gemini CLI** | `~/.gemini/settings.json` (`mcpServers` key) — `midas init` writes it |
| **Cline** | `cline_mcp_settings.json` in VS Code global storage — `midas init` writes it |
| **Zed** | `settings.json` → `context_servers` — `midas init` writes it |
| **Anything else** | point it at command `midas-mcp` |
| **No Python** | `npx -y midas-memory-mcp` — the [TypeScript port](packages/midas-ts) (experimental: no semantic embeddings yet) |
| **No Python** | `npx -y midas-memory-mcp` — the [TypeScript port](packages/midas-ts) (experimental; semantic embeddings via optional `@huggingface/transformers`) |

Override per client with env: **`MIDAS_MCP_DB`** (default `~/.midas/memory.sqlite3`; `:memory:` = ephemeral)
· `MIDAS_MCP_MAX_RECORDS` · `MIDAS_MCP_MIN_IMPORTANCE` · `MIDAS_MCP_NAMESPACE`.
Expand Down Expand Up @@ -186,15 +201,20 @@ server runs in. Or scope it manually per project/agent/user with `MIDAS_MCP_NAME
<summary><b>All tools & env knobs</b></summary>

**Tools:** `remember`, `capture` (policy-gated auto-store), `recall` (source-traceable), `build_context`
(compact, dated, today-anchored prompt block), `memory_state` (current project state), `memory_diff`
(what changed since), `check_memory_use` (guard), `memory_policy`, `maintain` (dedup + forgetting, returns
a deletion audit), `stats`, `forget` (chain-safe), `forget_matching` (topic-level erasure, dry-run by
default), `forget_all`. Prompts: `memory_session`, `distill`.
(compact, dated, today-anchored prompt block), `resume` (the one-call session-onboarding pack: pinned +
state + changes + open loops + conflicts), `memory_state` (current project state), `memory_diff`
(what changed since), `memory_conflicts` (live beliefs that contradict each other, ranked),
`open_loops` / `remember_commitment` / `close_loop` (promised work that survives sessions),
`check_memory_use` (guard), `memory_policy`, `maintain` (TTL + dedup + forgetting, returns a deletion
audit), `stats`, `forget` (chain-safe), `forget_matching` (topic-level erasure, dry-run by default),
`forget_all`. Prompts: `memory_session`, `distill`.

**Env:** `MIDAS_MCP_DB` · `MIDAS_MCP_EMBEDDER` (`local` / `hashing` / `multilingual` / any fastembed id) ·
`MIDAS_MCP_MAX_RECORDS` · `MIDAS_MCP_MIN_IMPORTANCE` · `MIDAS_MCP_NAMESPACE` (`=auto` → per-project scope) · `MIDAS_MCP_ANN=1` (sub-linear
IVF for huge stores) · `MIDAS_MCP_SUPERSEDE` · `MIDAS_MCP_NLI=1` (NLI-gated revision) ·
`MIDAS_MCP_AUTO_MAINTAIN=<min>` (idle-time upkeep) · `MIDAS_MCP_PINNED` (pin standing directives).
`MIDAS_MCP_AUTO_MAINTAIN=<min>` (idle-time upkeep) · `MIDAS_MCP_PINNED` (pin standing directives) ·
`MIDAS_MCP_TTL` (per-kind retention, e.g. `chat=30,note=90`) · `MIDAS_MCP_TOKEN` (HTTP bearer auth) ·
`MIDAS_MCP_KEY` (SQLCipher encryption at rest — `pip install "midas-memory[encrypted]"`).

</details>

Expand Down Expand Up @@ -260,10 +280,23 @@ midas inspect --db ~/.midas/memory.sqlite3 # opens http://localhost:7777
# before install: python -m midas.inspector --db <your.sqlite3> --embedder hashing
```

- **Browse + search** every memory (verbatim, with provenance + source).
- **Overview** — counts, attributability, a 30-day activity chart, and kind/provenance/recency breakdowns,
each kind and provenance color-coded *consistently across every view* (a fixed categorical palette,
validated for colorblind-safe contrast in both themes — never color-only, every value keeps its label).
- **Browse + search** every memory (verbatim, with provenance + source), filterable by kind, provenance,
and sort order.
- **Belief history + time-travel** — what you believed, what it superseded, and when.
- **Project state** (decisions / bugs / forbidden) and **what changed** since a date.
- **Governance** — would memory authorize an action, and why (the audit trail); **forget** with a receipt.
- **Conflicts** and **Open loops** — the same control-plane views from `memory_conflicts`/`open_loops`,
with one-click resolve/close from the UI.
- **Audit log** — the hash chain's verification status and its most recent entries.
- Light + dark themes (a real second theme, not an inverted dark one), keyboard shortcuts (**⌘K** to jump
anywhere or search, **/** to focus search), and a responsive layout down to phone width.

And every mutation (write / revise / forget) appends to a **tamper-evident, hash-chained audit log**
inside the store — hashes only, never content. `midas audit` shows it; `midas audit --json` verifies
the whole chain and reports the first broken entry if anyone rewrote history.

No LLM, no account, runs on your file. The thing a black-box memory can't show.

Expand Down Expand Up @@ -319,5 +352,7 @@ python -m eval.continuity #

Local-first: every memory lives in a SQLite file on your machine, recall returns the exact stored text,
and capture/recall/forget make **no network calls**. No account, API key, or telemetry. The only outbound
traffic is a one-time embedding-model download (for the `local` backend) and the package install. Full
details in [`PRIVACY.md`](PRIVACY.md) · [Apache-2.0](LICENSE).
traffic is a one-time embedding-model download (for the `local` backend) and the package install.
Optional **encryption at rest**: set `MIDAS_MCP_KEY` with the `[encrypted]` extra and the store is a
SQLCipher database — unreadable without the key (and Midas fails closed rather than silently writing
plaintext). Full details in [`PRIVACY.md`](PRIVACY.md) · [Apache-2.0](LICENSE).
3 changes: 3 additions & 0 deletions eval/benches.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,9 @@ def run(verbose: bool = True) -> dict:
("Agent Continuity", "action_safety", cont["action_safety"], "→1.00"),
("", "decision_adherence", cont["decision_adherence"], "→1.00"),
("", "repeated_mistake", cont["repeated_mistake"], "→1.00"),
("", "resume_fidelity", cont["resume_fidelity"], "→1.00"),
("", "conflict_detection", cont["conflict_detection"], "→1.00"),
("", "conflict_precision", cont["conflict_precision"], "→1.00"),
("Memory-Safety", "ASR (attack-success)", safe["ASR"], "→0.00"),
("", "benign_pass", safe["benign_pass"], "→1.00"),
("Coding-agent", "decision_currency", code["decision_currency"], "→1.00"),
Expand Down
Loading
Loading