semantic-memory-mcp is a local-first Model Context Protocol server for the
semantic-memory Rust library. It gives MCP clients
persistent semantic search, witnessed retrieval, durable receipts, governed
authority decisions, graph and lifecycle tools, and optional claim-ledger trust
enrichment over a store that remains on the operator's machine.
The default build uses SQLite/FTS5, the usearch vector backend, and an in-process Candle embedder. Ollama is an alternative embedder. The first Candle run downloads the configured Hugging Face model; after that, normal search and storage do not require a hosted database or API key.
The current Rust source and Cargo.toml are authoritative. In particular:
- The binary serves MCP over stdio.
--http-portcan add a loopback-only warm HTTP surface;--http-onlydisables stdio. leanandstandardare aliases in behavior and expose four governed, read-only tools.stableexposes a bounded 10-tool read-only surface.agentexposes a bounded 12-tool surface; its only write is governed fact capture (sm_add_fact).fullexposes every tool compiled into that build. Its size can change with feature selection, so this README does not freeze a full-profile tool count.- MCP
tools/listis the source of truth for the tools available in a specific binary, profile, and build. - The
fullCargo feature is the default and is currently an alias forsearch; it is unrelated to the runtime--tool-profile fullswitch.
Run without installing:
npx -y @recursiveintell/semantic-memory-mcp --memory-dir ~/.local/share/semantic-memory --tool-profile agentOr add to your MCP client config:
{
"mcpServers": {
"semantic-memory": {
"command": "npx",
"args": ["-y", "@recursiveintell/semantic-memory-mcp", "--memory-dir", "~/.local/share/semantic-memory", "--tool-profile", "agent"]
}
}
}Install the published package from crates.io:
cargo install semantic-memory-mcp --locked --version '=0.5.6'Then add to your MCP client config:
{
"mcpServers": {
"semantic-memory": {
"command": "semantic-memory-mcp",
"args": ["--memory-dir", "~/.local/share/semantic-memory", "--tool-profile", "agent"]
}
}
}From a checkout with the sibling path dependencies present:
cargo build --release
./target/release/semantic-memory-mcp \
--memory-dir "$HOME/.local/share/semantic-memory" \
--tool-profile agentdocker run -i --rm \
-v "$HOME/.local/share/semantic-memory:/data" \
ghcr.io/recursiveintell/semantic-memory-mcp:latest \
--memory-dir /data --tool-profile agentDocker compose:
services:
semantic-memory-mcp:
image: ghcr.io/recursiveintell/semantic-memory-mcp:latest
volumes:
- ~/.local/share/semantic-memory:/data
command: --memory-dir /data --tool-profile agent
stdin_open: trueThe Candle model defaults to nomic-ai/nomic-embed-text-v1.5 through the
nomic-embed-text alias. To use Ollama instead:
ollama pull nomic-embed-text
semantic-memory-mcp \
--memory-dir "$HOME/.local/share/semantic-memory" \
--embedder ollama \
--embedding-url http://localhost:11434 \
--tool-profile agent--memory-dir names a directory, not a database file. The store creates
memory.db and its index sidecars below that directory.
The options below match src/main.rs and the generated --help surface.
| Option | Required/default | Meaning |
|---|---|---|
--memory-dir <MEMORY_DIR> |
required | Store directory, created when absent. |
--embedder <EMBEDDER> |
candle |
candle, ollama, or the test-only mock backend. |
--embedding-url <EMBEDDING_URL> |
Ollama default http://localhost:11434 |
Used only by the Ollama backend. |
--embedding-model <EMBEDDING_MODEL> |
nomic-embed-text |
Candle Hugging Face model ID/alias or Ollama model name. |
--embedding-dims <EMBEDDING_DIMS> |
768 |
Embedding dimensions. Must match the chosen model/store. |
--http-port <HTTP_PORT> |
unset | Start the warm HTTP server on 127.0.0.1:<port> alongside MCP. |
--http-only |
false | Skip stdio MCP and keep only HTTP running. Requires --http-port for a useful process. |
--turbo-quant |
false | Request TurboQuant candidate generation with exact f32 reranking. The local full feature must be active for the bridge wiring to run. |
--turbo-quant-bits <TURBO_QUANT_BITS> |
codec default 8 |
Polar angle bits, used only with --turbo-quant. |
--turbo-quant-projections <TURBO_QUANT_PROJECTIONS> |
codec default 16 |
QJL projection count, used only with --turbo-quant. |
--tool-profile <TOOL_PROFILE> |
lean |
stable, lean, standard, agent, or full. Unknown values are rejected by typed CLI parsing before the store is opened. |
-h, --help |
— | Print generated help. |
The typed profile manifest in src/profile.rs supplies profile names, bounded
allowlists, effect classes, and HTTP effect capabilities. Use tools/list when
automating against a deployed binary.
These autonomous profiles expose exactly:
sm_search_witnessedsm_replay_searchsm_decide_assertion_authoritysm_decide_action_authority
They do not expose raw search, mutation, maintenance, import, or administration.
This bounded read-only profile exposes ten tools: ordinary and witnessed search, store statistics and namespace discovery, direct fact/neighbor lookup, graph paths, conversation retrieval, and separate assertion/action authority decisions. It excludes writes, imports, lifecycle administration, and maintenance.
The daily coding-agent profile exposes these 12 tools:
sm_add_fact sm_decide_action_authority
sm_decide_assertion_authority
sm_get_fact sm_get_fact_neighbors
sm_get_search_receipt sm_graph_path
sm_list_namespaces sm_replay_search
sm_search_conversations sm_search_witnessed
sm_stats
sm_add_fact is the only write. It remains governed and fails closed without a
trusted authority issuer. The profile excludes deletion, raw/unwitnessed
search, imports, lifecycle administration, reconciliation, vacuuming, and
re-embedding. Use lean for the smallest autonomous recall/authority surface;
use full only for explicitly approved operator work.
Some earlier deployments started without an in-process operator issuer. That
fail-closed configuration can look like an overly aggressive security policy:
even ordinary local sm_add_fact requests are rejected with no trusted issuer
instead of being appended. The supported recovery is not to remove all
admission controls. Start the MCP service with a private operator-token file,
which creates an in-process issuer only for its local governed capture path:
install -d -m 700 "$HOME/.config/semantic-memory"
umask 077
openssl rand -hex 32 > "$HOME/.config/semantic-memory/operator-authority.token"
semantic-memory-mcp \
--operator-authority-token-file "$HOME/.config/semantic-memory/operator-authority.token" \
--tool-profile agentUse a non-empty token without whitespace and keep the file readable only by the
service owner. The token is read at startup and consumed to construct the
in-process issuer; it is not sent as an MCP tool parameter. When the process
starts, require the operator authority: enabled log line, then append one
ordinary fact with a non-blank idempotency key and read it back by ID. Removing
the flag and restarting is the rollback: supported writes again fail closed.
With the issuer present, ordinary public or internal durable facts without
source or evidence references and without entity extraction can be appended
through sm_add_fact and return a governed receipt. This does not admit
unresolved provenance, confidential or restricted data, ephemeral_inference,
entity extraction, deletion, or other operator tools. Keep the MCP process on
its normal local/loopback transport boundary and retain transport authentication
where applicable; write admission is not network-access policy.
This is the operator profile. It exposes every tool registered by the compiled
router, including mutating, destructive, experimental, and maintenance tools.
Tool annotations describe read-only, idempotent, and destructive intent, but an
MCP client must still enforce its own approval policy. Inspect tools/list
before granting this profile to an autonomous process.
Tools exposed by the full router, grouped by function. lean/standard,
stable, and agent use the bounded manifests above; full exposes the
compiled router. Run tools/list against the deployed binary before automating
against any profile.
| Tool | Description | Profile |
|---|---|---|
sm_search |
Hybrid BM25+vector search with RRF fusion | full |
sm_search_witnessed |
Mandatory witnessed retrieval — bypasses cache, requires durable receipt, replay_mode opt-in | lean+ |
sm_search_with_routing |
Adaptive search: profiles query, routes to stages, optional factor graph + community grouping | full |
sm_search_proof_debt |
Search with trust-index gating and proof-debt budget | full |
sm_search_as_of |
Bitemporal search — facts valid at a specific date | full |
sm_search_conversations |
Hybrid search over stored conversation messages | agent+ |
| sm_route_query | Profile a query and get adaptive routing decision | full |
| sm_get_search_receipt | Load a durable search receipt by ID | agent+ |
| sm_replay_search | Replay a search using stored (opt-in) inputs | lean+ |
| sm_replay_search_receipt | Replay with caller-supplied query — compares results | full |
| sm_benchmark_trust | Benchmark trust quality distribution across searches | full |
| sm_get_routing_policy | Return current RL routing policy weights | full |
| Tool | Description | Profile |
|---|---|---|
sm_add_fact |
Add a fact — embedded, FTS-indexed, optional entity extraction | full |
sm_get_fact |
Fetch one fact by ID with full metadata | agent+ |
sm_get_fact_neighbors |
Fetch a fact plus its graph neighbors with content | agent+ |
sm_update_fact |
Update fact content in-place — re-embeds, updates indexes | full |
sm_delete_fact |
Hard delete — governed, irreversible | full |
sm_supersede_fact |
Create replacement + link via supersedes edge (preferred over delete) | full |
sm_consolidate_facts |
Merge two near-duplicate facts into one | full |
sm_list_facts |
Enumerate facts in a namespace (paginated) | full |
sm_list_namespaces |
List all namespaces with fact counts | agent+ |
sm_delete_namespace |
Permanently delete all memory in a namespace | full |
sm_set_provenance |
Set confidence (0.0–1.0) with support count | full |
sm_ingest_document |
Ingest with auto-chunking — each chunk embedded + indexed | full |
| Tool | Description | Profile |
|---|---|---|
sm_add_graph_edge |
Add typed edge: Semantic, Temporal, Causal, or Entity | full |
sm_list_graph_edges |
List edges for a node or all edges | full |
sm_invalidate_graph_edge |
Append-only invalidation — never deletes | full |
sm_graph_path |
BFS shortest path between two nodes | agent+ |
sm_community |
Leiden-inspired community detection with contradiction scanning | full |
sm_topology |
Topological void detection — Betti numbers | full |
sm_factor_graph |
Belief propagation over all 4 edge types | full |
sm_decoder_analyze |
Contradiction detection + belief propagation refinement | full |
sm_detect_contradictions |
Content-based contradiction signals (no pre-asserted edges needed) | full |
sm_discord_search |
Second-order graph traversal from direct search results | full |
sm_subgraph_prune |
Access-frequency-based subgraph pruning (dry-run default) | full |
| Tool | Description | Profile |
|---|---|---|
sm_create_claim |
Create typed claim from a fact with source-spanned provenance | full |
sm_add_evidence |
Add EvidenceBundle supporting a claim | full |
sm_judge_support |
Judge claim support: supported / unsupported / contested / heuristic_only | full |
sm_verify_claim |
Risk-class-based verification (low→critical, falsification for high+) | full |
sm_decide_assertion_authority |
Governed assertion decision receipt — purpose-isolated | lean+ |
sm_decide_action_authority |
Governed action decision receipt — purpose-isolated | lean+ |
sm_query_claim_versions |
Bitemporal claim projection queries | full |
sm_query_relation_versions |
Bitemporal relation projection queries | full |
sm_query_episodes |
Episode projection queries | full |
sm_query_entity_aliases |
Entity alias projection queries | full |
sm_query_evidence_refs |
Evidence reference projection queries | full |
sm_compact_claim_ledger |
Verified hash-chained claim ledger rotation | full |
| Tool | Description | Profile |
|---|---|---|
sm_stats |
Store statistics: counts, DB size, embedding model | agent+ |
sm_run_lifecycle |
Syndrome detection, subtraction candidates, compression assessment | full |
sm_reconcile |
Integrity actions: ReportOnly / RebuildFts / ReEmbed | full |
sm_vacuum |
SQLite VACUUM — compact and defragment | full |
sm_reembed_all |
Re-embed all facts after embedding model change | full |
sm_embeddings_are_dirty |
Check if embeddings need regeneration | full |
| Tool | Description | Profile |
|---|---|---|
sm_import_envelope |
Atomic bulk import with provenance — single transaction | full |
sm_import_status |
Check envelope import status (idempotency) | full |
sm_list_imports |
List recent imports, optionally filtered by namespace | full |
| Tool | Description | Profile |
|---|---|---|
sm_parse_json |
Extract JSON from raw LLM output (handles think blocks, fences) | full |
sm_parse_json_value |
Parse as untyped serde_json::Value | full |
sm_parse_choice |
Parse a choice from valid options list | full |
sm_parse_number |
Parse a number from raw LLM output | full |
sm_parse_string_list |
Parse string list from bullet/comma/JSON formats | full |
sm_repair_json |
Repair common LLM JSON errors | full |
sm_strip_think_tags |
Strip </think> blocks from text |
full |
sm_record_outcome |
Record RL routing feedback label for policy training | full |
With --http-port, the loopback server exposes:
GET /health — health check (all profiles)
GET /verify-integrity — integrity check across all surfaces
POST /search — hybrid search
POST /search-routed — adaptive search with routing
POST /rerank — exact f32 cosine rerank
POST /stats — store statistics
POST /add — add fact
POST /record-outcome — RL feedback
POST /discord — discord (second-order) search
POST /maintenance/check — maintenance health
POST /maintenance/vacuum — SQLite vacuum
POST /maintenance/reembed — re-embed all
POST /maintenance/reconcile — reconcile integrity
POST /maintenance/rebuild-hnsw — rebuild HNSW index
POST /maintenance/compact-hnsw — compact HNSW index
Lean/standard/agent profiles expose only /health. Full exposes all routes. All non-health endpoints require Bearer token auth.
sm_search_witnessed is the safe autonomous retrieval surface. It bypasses the
cache, requires a durable receipt, defaults to current state, and only returns
persisted facts whose source provenance can be hydrated honestly. Its
retrieval_mode is hybrid, fts_only, or vector_only.
V35 complete replay is privacy-sensitive and opt-in. The default
replay_mode: "no_replay" stores receipt digests and result evidence without
retaining the query/filter inputs needed for complete replay. Set
replay_mode: "store_inputs" on witnessed search only when that retention is
acceptable, then call sm_replay_search with the original receipt ID.
sm_replay_search_receipt remains a full-profile alternative that requires the
caller to resupply the query and filters.
Recall authority never implies permission to assert a result as true or to act
on it. sm_decide_assertion_authority and sm_decide_action_authority make
separate, fixed-purpose decisions from caller, subject, audience, namespace
scope, and an optional delegation/elevation lease. They return a typed decision
receipt and intentionally omit memory content; neither tool performs the
assertion or action.
The semantic store and the trust ledger have different jobs:
memory.db, FTS5, and vector/sparse indexes hold searchable memory.- Before first compaction,
claim_ledger.jsonlis the hash-chained trust authority. - After compaction, an atomically selected, digest-verified snapshot plus retained JSONL tail represents the same ledger history. The snapshot is a checkpoint, never an independent truth store.
- A process-local
ClaimTrustIndexis derived from the verified snapshot and tail (or the legacy JSONL) at startup. It is an acceleration structure and is never persisted as authority.
Search trust enrichment uses six quality states:
| State | Meaning | Default proof debt |
|---|---|---|
supported |
Recorded evidence supports the linked claim. | none |
partially_supported |
Recorded evidence supports only part of it. | none in the current mapper |
unsupported |
The judgment rejects support. | missing source basis |
contradicted |
Contradictory evidence has been recorded. | missing source basis + missing reproduction |
heuristic_only |
The judgment is heuristic rather than evidentiary. | missing source basis |
persisted_unjudged |
The fact has no linked judgment, or claim integration is absent. | missing source basis only when a linked claim exists; no claim means no claim debt to score |
sm_search_proof_debt exposes debt-aware retrieval and a budget gate;
sm_benchmark_trust reports the distribution of the six states. Proof debt is
an obligation signal, not a replacement for source inspection.
If the legacy ledger, active manifest, snapshot, retained tail, or compaction receipt fails verification, the server disables claim trust enrichment and refuses ledger append/compaction. Ordinary semantic storage and search remain available. Results report trust enrichment as disabled where that path can surface it; corruption does not promote an unverified ledger or erase the semantic database.
sm_compact_claim_ledger is a destructive-annotated, full-profile,
claim-integration tool. It defaults to a dry run:
{
"dry_run": true,
"max_entries": 10000,
"max_bytes": 16777216,
"retain_tail_entries": 256,
"max_backups": 3
}No publication occurs unless a threshold is exceeded and dry_run is
explicitly false. A real compaction writes and fsyncs a temporary generation
containing snapshot.json, tail.jsonl, and receipt.json, renames that
generation into place, then atomically replaces
claim_ledger.active_compaction.json. That manifest rename is the publication
boundary: startup ignores incomplete temporary generations and accepts only the
manifest-selected generation after digest verification.
The production search path is implemented by semantic-memory:
- Embed or tokenize the query.
- Retrieve FTS5/BM25 and usearch vector candidates.
- When configured and represented by the active embedder, retrieve V36 sparse dot-product candidates from SQLite.
- Fuse ranks with RRF and apply temporal/provenance policy.
- Filter superseded heads for the normal MCP search surfaces.
- Persist receipts and, for witnessed search, hydrate source provenance, authority state, and optional claim-ledger trust.
V36 sparse storage and ranking are inherited from semantic-memory; this MCP
crate does not define a separate sparse feature or CLI switch. The default
SearchConfig has sparse_weight = 0, so sparse retrieval is dormant unless a
library-level configuration enables it. The active embedder must also provide a
sparse representation, or explicit dense-to-sparse derivation must be enabled.
Advanced full-profile tools can additionally route queries, explain ranking, traverse stored graph edges, detect contradictions, run factor-graph analysis, inspect topology/communities, and perform lifecycle or maintenance work. Those surfaces are not implied by the four-tool autonomous profile.
These are the exact local features declared in Cargo.toml:
| Feature | Default? | What it enables |
|---|---|---|
default |
yes | full |
full |
via default | Alias for search; also activates local cfg(feature = "full") wiring such as TurboQuant candidate selection. |
search |
via full |
The supported composed router build: usearch, Candle, provenance, temporal, multiscale, discord, decoder, subtraction, compression governor, routing, admin ops, late interaction, TurboQuant codec, RL routing, plus the local integration features below. |
integration |
via search |
Forwards semantic-memory/integration. |
subgraph-pruning |
via search |
Forwards semantic-memory/subgraph-pruning and enables sm_subgraph_prune. |
candle-embedder |
via search |
Forwards the in-process Candle backend. |
claim-integration |
via search |
Adds the optional claim-ledger dependency and claim/trust/compaction tools. |
llm-parser |
via search |
Adds the optional llm-output-parser dependency and parser tools. |
orchestration |
via search |
Adds knowledge-runtime and provenance/temporal orchestration tools. |
hnsw |
no | Forwards the alternative semantic-memory/hnsw backend and enables compact/rebuild HNSW endpoint code where gated. |
cargo build --no-default-features --features search compiles the composed
search feature without setting the local full cfg. It is not a minimal tool
surface. Builds assembled from narrower individual features are feature-gated
development configurations, not the documented production default.
Production-wired in the default build:
- stdio MCP, runtime tool-profile filtering, SQLite/FTS5, usearch, Candle and Ollama embedders;
- witnessed retrieval, V35 opt-in replay, governed assertion/action decisions, receipts, graph storage/traversal, claim-ledger verification and compaction;
- the composed
semantic-memoryrouting, provenance, temporal, decoder, lifecycle, orchestration, parser, and admin capabilities exposed by the full router.
Opt-in, feature-gated, or operationally experimental:
- the loopback HTTP server is an auxiliary API, not MCP and not authenticated;
mockembeddings are for tests;hnswis an optional alternative to the default usearch backend;- TurboQuant requires
--turbo-quantand the localfullcfg wiring; - LLM reranking, entity extraction, and community summaries call a local Ollama service and are opt-in per operation;
- V36 sparse retrieval is inherited and disabled by the default search weight;
- broad maintenance, deletion, import, training-feedback, and lifecycle tools are operator-only even when compiled.
With --http-port, the process binds only to 127.0.0.1. Current routes are:
GET /health
GET /verify-integrity
POST /search
POST /search-routed
POST /rerank
POST /stats
POST /add
POST /record-outcome
POST /discord
POST /maintenance/check
POST /maintenance/vacuum
POST /maintenance/reembed
POST /maintenance/reconcile
POST /maintenance/rebuild-hnsw
POST /maintenance/compact-hnsw
The HTTP sidecar applies the selected profile below transport. Lean, standard,
and agent expose only /health; the explicit full operator profile exposes the
authenticated non-health surface. All non-health requests require a valid bearer
token. Mutation handlers without a trusted authority issuer fail closed. All
requests still require loopback Host/Origin validation.
First-class packages live in integrations/:
- Hermes plugin — validates inputs and invokes the
current
hermes mcp add/list/test/configureCLI workflow without reimplementing any semantic-memory tools. - Claude Code plugin — manifest, plugin-scoped MCP launcher, semantic-memory skill, and useful commands.
- Codex integration — open Agent Skill layout plus stdio MCP installation/config examples.
- Install/test matrix — side-by-side setup and smoke checks.
- Memory contents, sources, conversation messages, replay inputs, and claim evidence may be sensitive. Protect the entire memory directory with OS-level permissions and backups appropriate to its data classification.
- Candle's model download contacts Hugging Face on first use. Ollama mode sends text to the configured Ollama URL; a remote URL moves content off-host.
store_inputsretains query/filter material for complete replay. It is off by default for privacy.- The claim ledger is tamper-evident, not encrypted. Verification detects corruption; it does not stop a party with filesystem access from reading it.
- The
fullprofile and HTTP maintenance routes include mutation, permanent deletion, model-feedback, import, vacuum, and rebuild operations. Grant them only to an operator context with explicit approval controls. agentis the recommended profile for trusted coding agents. It is read-only until a trusted authority issuer is injected. Useleanfor autonomous read-only recall and authority decisions.- Do not place secrets in facts, sources, metadata, replay inputs, plugin config, or command-line arguments. Process lists and logs may expose arguments.
Use the current CLI rather than hand-editing legacy YAML as the primary path:
hermes mcp add semantic_memory \
--command semantic-memory-mcp \
--args --memory-dir "$HOME/.local/share/semantic-memory" --tool-profile agent
hermes mcp list
hermes mcp test semantic_memory
hermes mcp configure semantic_memory--args consumes the remaining arguments, so it must come last. The configure
step is interactive and controls which server-native tools Hermes exposes. The
packaged Hermes plugin provides guarded wrappers for these commands; see its
README.
Test the local plugin directly:
SEMANTIC_MEMORY_MCP_BIN="$(command -v semantic-memory-mcp)" \
SEMANTIC_MEMORY_DIR="$HOME/.local/share/semantic-memory" \
SEMANTIC_MEMORY_TOOL_PROFILE=agent \
claude --plugin-dir ./integrations/claude-pluginInside Claude Code, run /mcp, then
/semantic-memory:semantic-memory-status. Use claude --debug --plugin-dir ./integrations/claude-plugin for startup diagnostics and claude plugin validate ./integrations/claude-plugin when supported by the installed Claude
Code version.
Install the stdio server with the CLI:
codex mcp add semantic_memory -- \
semantic-memory-mcp \
--memory-dir "$HOME/.local/share/semantic-memory" \
--tool-profile agent
codex mcp listCopy or symlink integrations/codex/.agents/skills/semantic-memory into the
repository's .agents/skills/, or keep the supplied structure at the project
root. Codex discovers skills from .agents/skills between the working directory
and repository root. See the Codex integration README for a TOML alternative.
cargo fmt --check
cargo check
cargo testIntegration asset validation is read-only:
python3 integrations/tests/validate_integrations.pyApache-2.0. See LICENSE.