Reviewed: 2026-09-15 · Code-grounded.
The Knowledge Base is a fully offline document store with hybrid retrieval. An operator uploads documents; the node extracts their text, chunks them on header boundaries, embeds each chunk with a local embedding model, and indexes everything into local SQLite with selective encryption — source document blobs and display names are encrypted at rest, while the extracted chunk text and its FTS search index are stored unencrypted locally. Retrieval fuses a lexical arm (SQLite FTS5 / BM25) and a semantic arm (vector cosine similarity) with Reciprocal Rank Fusion, optionally rescoring with a local cross-encoder reranker, and exposes the result to agents as a tool.
Provider-locality invariant (enforced default). Ingestion and embedding stay on-device by default (see KnowledgeBase:EmbeddingProviderName), and — as of the MED-004 hardening — the read-only knowledge tools (search_knowledge_base, read_document, read_surrounding_chunks) are offered only to node-local models (llama.cpp / Ollama). A cloud-hosted model (Codex OAuth, Azure Foundry) is not offered these tools, so document text, chunks, and the query are never handed to a third-party provider through a tool call. An operator can waive this by setting KnowledgeBase:AllowCloudModelAccess=true (default false) — an explicit, documented opt-in that acknowledges knowledge-base content may then reach the cloud provider.
The gate keys on the EFFECTIVE model — the one that actually runs the turn after any agent/profile pin is applied — not the turn's active model. So an agent, orchestration participant, or spawned sub-agent pinned to a cloud model is withheld the knowledge tools even when the conversation's active model is local (each effective model's locality is resolved through the shared IModelCapabilityResolver; participants and pins are classified individually, spawned children through the same resolver keyed on the child's pin). Retrieved chunk text — and its attacker-controlled metadata (title, section, source) — is additionally returned to the model fenced as untrusted data: a contentTrust: "untrusted-document" flag plus per-response nonce-delimited begin/end markers (a random marker suffix the document body cannot forge), so an instruction embedded in a document, title, or heading cannot read as a system directive or break out of the fence. Nothing reaches a log.
Attachments and workspace file tools are gated by the same effective-model locality, with a visible notice. Conversation attachments (inlined in plain chat, or staged into the AgentHome workspace and read via the coder list_files/read_file/search_text tools) are node-local private data. When the effective model that runs a turn is cloud-hosted — including when an agent/profile pin substitutes a cloud model even though the user's active pick is local — and the operator has not opted in via KnowledgeBase:AllowCloudModelAccess, the attachments are withheld: not staged for the file tools, not inlined into the prompt, and the coder file tools are withheld from the offer. The user gets a visible turn notice naming the effective model. (The earlier "user chose the conversation model" rationale does not hold once a pin substitutes the model, so attachments are gated exactly like the Knowledge Base rather than waved through.) The opt-in restores them. Whatever content does reach the model is fenced as untrusted data (file name inside the fence) so it is never read as instructions; the attachment fence nonce is derived from a server secret (HKDF over the node key), not the client-visible conversation id, so a client cannot forge the fence's closing marker.
Config key. KnowledgeBase:AllowCloudModelAccess (default false) is the single opt-in governing all node-local private-data exposure to a cloud model: the knowledge tools, the coder workspace file tools, and conversation attachments. It is named under KnowledgeBase for continuity; set it to true to allow a cloud model to receive any of that node-local content.
| Concern | Project / path |
|---|---|
| Ingestion pipeline driver | XE-Local-AI-Engine.Client.Application/Services/Knowledge/Implementation/KnowledgeIngestionService.cs (IKnowledgeIngestionService) |
| Background ingestion worker + dispatcher | …/Services/Knowledge/Implementation/KnowledgeIngestionWorker.cs, KnowledgeIngestionDispatcher.cs |
| Header-boundary chunker | …/Services/Knowledge/Implementation/HeaderBoundaryChunkingService.cs (IChunkingService) |
| Chunk embedder | …/Services/Knowledge/Implementation/KnowledgeChunkEmbedder.cs (IKnowledgeChunkEmbedder) |
| Exact chunk-embedding reuse | …/Services/Knowledge/Implementation/KnowledgeChunkEmbeddingCache.cs, KnowledgeChunkEmbeddingReuseStore.cs |
| Embedding model resolver + prefixer | …/Services/Knowledge/Implementation/EmbeddingModelResolver.cs, KnowledgeEmbeddingPrefixer.cs |
| Hybrid search orchestrator | …/Services/Knowledge/Implementation/KnowledgeSearchService.cs (IKnowledgeSearchService) |
| Optional-stage policy | …/Services/Knowledge/AdaptiveRetrievalPolicy.cs |
| Lexical arm (FTS5) | …/Services/Knowledge/Implementation/FtsSearch.cs (IFtsSearch) |
| Semantic arm (cosine) + factory | …/Services/Knowledge/Implementation/ManagedCosineVectorSearch.cs, VectorSearchFactory.cs (IVectorSearch) |
| Encrypted document blob store | …/Services/Knowledge/Implementation/KnowledgeDocumentBlobStore.cs (IKnowledgeDocumentBlobStore) |
| Document text extraction | …/Services/DocumentIngestion/Implementation/DocumentTextExtractor.cs + Extraction/ |
| Registered-repository import | …/Services/Knowledge/Implementation/KnowledgeRepositoryImportService.cs |
| Scheduled stale-model reindex | …/Services/Knowledge/Implementation/KnowledgeScheduledModelReindexWorker.cs |
| Agent-facing tools | …/Services/Knowledge/Tools/Implementation/SearchKnowledgeBaseToolHandler.cs, ReadSurroundingChunksToolHandler.cs |
| SignalR notifier + hub | XE-Local-AI-Engine.Client/Hubs/KnowledgeIndexingNotifier.cs, KnowledgeBaseHub.cs |
| Local endpoints | XE-Local-AI-Engine.Client/Endpoints/Knowledge/V1/ |
| React feature | XE-Local-AI-Engine.Client.React/src/features/knowledge/ |
| Retrieval evaluation corpus/harness | XE-Local-AI-Engine.Client.Persistence.Tests/Knowledge/RetrievalEval/ |
These types hold a NodeChatDbContext and speak raw ADO from the application layer. That is a reasoned permanent exception to the store rule in Code conventions, not migration debt, and Architecture/DbContextUserAllowlist.txt records it per type. Three measured grounds: KnowledgeDocumentStatus is bound as a SQL parameter by half of them and is a shipped OpenAPI schema id and a SignalR parameter, so relocating the SQL would rename a generated client symbol; KnowledgeDocumentCatalogService.ResetStaleDocumentsToPendingAsync evaluates the embedding-model staleness policy between its SELECT and its UPDATEs, and KnowledgeDocumentBlobStore.UpdateRepositoryDocumentAsync holds its transaction open across an encrypted blob write and atomic rename, rolling database and file back together; and four of these interfaces are injected by host endpoints, which EndpointDependencyTests forbids for a persistence type. The two startup maintenance jobs (KnowledgeVectorNormalizationBackfillService, KnowledgeBlobOrphanSweeper) are listed with the other database-maintenance jobs instead.
KnowledgeIngestionService drives one document through the pipeline, advancing knowledge_documents.status at each transition and setting a content-free failure_reason on any step failure. Uploads are processed asynchronously by KnowledgeIngestionWorker so the upload endpoint returns as soon as the document is admitted to the queue and the operator watches status over the hub.
Upload ──▶ Uploaded ──▶ Extracting ──▶ Chunking ──▶ Embedding ──▶ Indexed
│ │ │ │ │
encrypted blob text extract header local model atomic index
(blob store) (.pdf/.docx/ boundaries batches write (FTS + vectors)
plaintext) chunker
- Extraction —
DocumentTextExtractordispatches by extension:.pdfand.docxhave dedicated readers (.docxvia the pure-managed Open XML SDK);.html/.htmuse the deterministicHtmlDocumentReader; Markdown, plaintext, structured text, logs, and supported source-code extensions usePlaintextDocumentReader. Markdown headings and fenced code blocks, HTML headings/paragraphs/lists/tables/code blocks, and code whitespace/symbol text are retained as ordered structural elements instead of being flattened prematurely. Extraction is bounded byDocumentExtractionLimits(a container format such as a.docxzip or.pdfcan expand well beyond its on-disk size, so a decompression-bomb guard caps the extracted size). - Chunking —
HeaderBoundaryChunkingServiceis an offline, deterministic chunker with no tokenizer or external package. It walks the document's ordered element stream: each header opens a new section (maintaining anH1 > H2heading trail via a level stack); paragraphs, tables, and code blocks accumulate into their section body, which is split into token- and character-bounded, overlapping chunks (MaxChunkTokens,MaxChunkChars,ChunkOverlapChars). Heading trails are prepended to the embedding input as deterministic document context. Each chunk preserves offsets plus content kind, language, symbol, page, and source-path provenance where the parser can supply them. The same document and versioned parser/chunker settings yield the same sections, chunks, hashes, and identifiers. Sizing is token-aware: a section body is cut at whichever bound is reached first — the per-chunk token budget (MaxChunkTokens, tightened when a resolved embedding context window is known to that window minus a safety reserve covering thesearch_document:intent prefix and the model's own special tokens, and floored so a tiny or misconfigured window cannot collapse chunking to near-single-character windows) or the hard character ceiling (MaxChunkChars) — always breaking at the last whitespace before the limit so no word is cut. Token-awareness is what keeps a chunk plus its heading prefix inside the embedding model's context window and stops a byte-pair tokenizer splitting a max-sized chunk mid-token. The per-window token budget is reduced by the section's heading-trail cost, so the embedded heading-prefixedContextualContentstays inside the window. Token counts come fromChunkTokenApproximation, a deterministic dependency-free approximation: weighted characters ÷ 4, with CJK and emoji weighted heavier so a token-dense script does not silently produce chunks several times the intended token size. The token budget can only tighten the effective size, never enlarge it past the character ceiling — plain ASCII prose therefore keeps the character ceiling as its binding bound and existing ASCII corpora need no reindex, while token-dense content and smaller embedder windows yield correspondingly smaller chunks. - Embedding —
KnowledgeChunkEmbedderreuses the node-local embedding resolution path: it resolves the configured embedding provider/model (KnowledgeBaseOptions.EmbeddingProviderName/EmbeddingModelName) and generates in batches (MaxEmbeddingBatchSize). After provider generation, ingestion and query shareKnowledgeEmbeddingVectorPolicy. A confidently resolvednomic-embed-text-v1.5defaults to the versioned Matryoshka recipe (full-vector population layer norm with epsilon1e-5→ first 512 components → L2 normalization); other models stay native. The provider's immutable installed-weight digest/revision is hashed into the resolved inventory fingerprint, so replacing weights under the same Ollama tag or GGUF name changes the canonical identity. The exact resolved model revision + transform/version + width is persisted and filtered during search, with dimension checked separately. When that identity and width are known before inference, ingestion reuses an exact vector keyed by contextual-content hash, parser version, chunker version, canonical vector identity, and dimension. A process-wide TTL/entry/byte-bounded RAM cache and in-flight coalescer first consult a short-lived scoped durable lookup of committed indexed rows; unknown-width providers retain the normal embed path rather than risk a false hit. - Indexing — the final
Indexedtransition is performed atomically byIKnowledgeIndexWriterso a document is never half-visible. All failure logging is exception-type-only — no chunk or document text reaches a log. - Admission control (backpressure) — the async queue is bounded (
KnowledgeIngestionDispatcher: a single-reader channel, capacity 256; the worker further caps concurrency with aSemaphoreSlim). Admission is non-blocking: when the queue is full, the upload/reindex endpoint returns 503 withRetry-After: 5instead of holding the request or growing the backlog, and admission is idempotent (a document already queued or in flight is not re-enqueued, so it is never processed twice concurrently). A 503 is a retryable busy signal, not data loss — the blob is already persisted, so a document whose ingestion a full queue previously rejected (persisted-but-unindexed) is re-enqueued on a later re-upload of the same file. Accept/reject counts and live queue depth are published on theXE.Nodemeter. The synchronous chat-attachment path has its own equivalent gate (DocumentExtractionAdmissionGate), also 503 +Retry-Afterwhen at capacity — see the busy-admission note in API & Hubs. - Worker lifecycle — the queue is an in-memory channel, so a document left mid-pipeline by a crash or hard stop would be stuck non-terminal forever: on start
KnowledgeIngestionWorkerresets every non-terminal document toPendingand re-dispatches it before draining new work. At shutdown it stops reading the queue and awaits the tracked in-flight documents within a bounded drain window (ShutdownDrainTimeoutSeconds) before disposal proceeds, so the scope factory and the semaphore are never disposed under a running document. Per-document work runs on a drain-deadline token cancelled only when that window elapses — ordinary operation never cancels a document mid-write, yet a hung document cannot block shutdown forever, and anything unfinished is left non-terminal and re-queued on the next start. A separate drain sweep re-admits documents stillPendingonce the admission queue empties, so uploads a full queue rejected are picked up as capacity frees; admission is idempotent, so nothing is ingested twice. - Scheduled index migration —
KnowledgeScheduledModelReindexWorkerperiodically discovers documents whose stored canonical vector identity, immutable model revision, parser version, or chunker version no longer matches the active pipeline and admits them to the same bounded background queue. Deterministic parser/chunker migrations still mark rows stale during a transient model-provider outage; model identity comparisons require a confident provider resolution — an installed model actually matched, never the configured-name fallback that an unreachable provider or an unmatched configured name produces. On a llama.cpp node the storedembedding_modelis a resolved GGUF name that never equals the plain configured name, so comparing against that fallback during an outage would flag, and reset, the entire indexed corpus instead of leaving it untouched. OnlyIndexedrows can be stale at all: a non-indexed row still carries the upload-time placeholder name, which is not a vector identity. It does not bypass normal queue pressure or atomic indexing.ScheduledModelReindexEnabledandScheduledModelReindexIntervalMinutescontrol the scheduler. BumpingKnowledgeIndexVersions.Parserre-queues every indexed document and re-embeds all of its chunks, because the parser version is part of the embedding reuse key.
Matryoshka rollback procedure. Set KnowledgeBase:EmbeddingVectorMode to Native, invoke the normal corpus reindex endpoint, and wait until the catalog reports no stale documents. Only then deploy a binary that predates canonical vector identities. Switching the binary first is unsafe because old model-name-only search code cannot distinguish the 512-wide transformed rows from native vectors. Switching either direction changes the canonical identity immediately, so old projections stay excluded/stale until the normal per-document or full-corpus reindex rebuilds them.
Every document, section, chunk, FTS row, vector candidate, and hydrated hit belongs to a canonical CollectionId. Missing collection input maps to DEFAULT for backward compatibility; explicit ids are trimmed, upper-cased, limited to 128 ASCII letters/digits plus -, _, and ., and applied independently to both lexical and dense retrieval. A DocumentId from another collection returns no results rather than weakening the collection boundary. This is a local namespace/isolation boundary, not a multi-tenant authorization system.
Development Mode can import a previously registered local Git repository through ImportKnowledgeRepositoryEndpoint. The request carries only the registered selected-folder id and collection id—never an arbitrary path. KnowledgeRepositoryImportService revalidates the registered folder as a local Git top-level, refuses reparse/symlink escapes, skips sensitive/generated directories and files, enforces per-file plus aggregate byte limits, revalidates the opened file before accepting its bytes, preserves only relative source paths, and sends eligible source files through the normal encrypted-blob/background-ingestion pipeline. Repository identity is collection + normalized source path: re-import updates a changed path in place, identical bytes at two paths keep distinct provenance, unchanged paths are not requeued, and a successfully completed scan purges repository documents whose paths were deleted or renamed. Uploads retain collection-scoped content-hash deduplication.
KnowledgeSearchService.SearchAsync is the retrieval heart. A KnowledgeSearchRequest carries the untrusted Query, a Limit, a normalized collection scope (defaulting to DEFAULT), an optional DocumentId scope, and an ExpandNeighbors flag. The flow:
- Embed the query with the current model, using the query-intent prefix (
KnowledgeEmbeddingPrefixer). - Two arms retrieve candidates in parallel:
- Lexical —
FtsSearchruns a BM25-rankedMATCHover weighted FTS5 fields for source path, heading path, symbol, and body content (the raw-SQL path; the query is escaped before it reaches FTS). Exact identifiers, file paths, and symbols therefore receive more weight than repeated body terms. Collection and optional document scopes are enforced in the same query through UNINDEXED identity columns. - Semantic — the collection- and model-scoped
IVectorSearch(ManagedCosineVectorSearch) scores chunk embeddings by cosine similarity within the active canonical model identity/dimension. Each arm fetches a candidate pool (CandidatePoolMultiplier× limit, floored atMinimumCandidatePool) so fusion has enough overlap material. The two arms are launched together, but only the lexical arm touches the request-scoped SQLite connection — the query-embedding arm talks solely to the embedding provider process — so overlapping them can never run two commands on that non-thread-safe connection at once. Every DB-bound read that consumes the embedding (the vector scan, hydration, neighbor expansion) runs sequentially afterwards on that one connection; truly concurrent execution of both DB arms would need a second connection/scope and is deliberately not taken. The win is overlapping the embedding round trip, typically the dominant latency, with the FTS query. The vector arm is filtered by the same resolved model name the query was embedded with, so query vectors are only ever compared against chunk vectors built by that identical model — a later same-dimension model swap changes the name and excludes the now-incompatible old vectors instead of silently mis-comparing them.
- Lexical —
- Fuse the two ranked lists with Reciprocal Rank Fusion (
IRankingFusionService). The fused set is the union of every chunk id either arm returned. Two strategies exist.Rrfis the classic score-agnostic fusion that uses rank position only — the graceful fallback and the comparison baseline.ScoreAwaremin-max normalizes each arm's scores within that arm, so incomparable scales (BM25 magnitude vs cosine) never blend directly, and uses the normalized value to tilt the RRF contribution multiplicatively by up to a configured weight, so a rank whose arm score sits far above its arm's floor outranks an equally-ranked but marginal competitor.ScoreAwaredegrades to pure RRF for any arm whose scores are empty, single, constant, or non-finite, because normalization carries no signal there. Callers orient incomparable raw scores before fusion so that higher always means more relevant — FTS5 BM25, which is more-negative-for-stronger, is negated. - Conditionally rerank the fused candidate pool with a local cross-encoder reranker (
IRerankerClient, modelKnowledgeBaseOptions.RerankerModelName) before the top-limitcut. WithAdaptiveRerankingEnabled(default), the deterministic policy skips the optional model when both retrieval arms agree on the top chunk, the candidate set is too small, or at least 80% ofRetrievalLatencyBudgetMilliseconds(default 500 ms) has elapsed; ambiguous candidates still rerank. The same budget, measured from search start, is the hard deadline for the rerank stage. Search never spawns the reranker:LlamaServerRerankerClientonly uses a process that is already running, so a cold reranker returns null at once (degrade reasoncold), the search returns fusion order, andKnowledgeSearchServiceasksIKnowledgeModelPrewarmerfor a background warm; start-up time is therefore never inside the budget. Scoring is abandoned, not cancelled, at the deadline: llama-server keeps scoring a cancelled request anyway, so the HTTP call finishes in the background under the client's own 5–30 s timeout, and while it is outstanding every new rerank for that model falls back at once (degrade reasonbusy). Deadline expiry degrades to fusion order; caller cancellation still propagates. Disabling the adaptive gate restores always-attempt-rerank behavior for controlled benchmarks, but not an unbounded wait. Every hit carries aScoreKindnaming the scale of itsScore, one kind per result because reranking is all-or-nothing:Fusionis the fused RRF score (roughly 0.01–0.06),Rerankthe raw cross-encoder relevance, set only when reranking succeeds (unbounded and model-specific: bge roughly −11…+7, qwen3 0…1). Every fallback (reranker off, gate skip, null or count-mismatched response, deadline expiry) staysFusion. Neither is a probability, and scores compare only within one result. The REST hit (scoreKind:"Fusion"/"Rerank"), the agent tool JSON and persisted chat sources carry it; a chat source saved before it existed loads with a null kind. - Hydrate the selected chunks over the raw-SQL path and, when
ExpandNeighborsis set, expand each hit with its surrounding neighbor chunks (NeighborWindow = 1each side).
Pre-warm. KnowledgeModelPrewarmer (singleton) ensures the configured reranker and, on the llama.cpp provider, the resolved embedder through EnsureRunningAsync. InvocationRunner requests it once per chat or agent turn whose offer contains search_knowledge_base (with AgentToolsEnabled on), after the turn's own model is ready so the two launches never race; a search whose rerank came back empty, malformed or past the retrieval budget requests it again, and that warm's reuse path re-probes a live reranker's liveness, so one that accepts and hangs is torn down and respawned. Requests coalesce into one warm in flight; after a failed or refused warm, requests are ignored for 30 s so repeated searches do not repeat the capacity probe. A successful warm sets no cooldown: re-ensuring a running process is a cheap reuse that also refreshes its idle timestamp, since pooled processes are still reaped after the idle TTL.
Graceful degradation is a contract: if the embedding model or the reranker is unavailable, search degrades (lexical-only / fusion-order) rather than failing. Each hit carries collection, document, chunk, relative source path, heading/section, page, offsets, content kind, language, and symbol provenance where available. A hit's Title/Section are derived from the non-sensitive heading_path/storage_path, so a result never exposes the encrypted original file name.
Client.Persistence.Tests/Knowledge/RetrievalEval/ runs the real ingestion, SQLite FTS/vector, fusion, hydration, and optional reranker path against deterministic embeddings. Its representative corpus covers English and German semantic questions, exact code identifiers/paths, lexical distractors, long-document boundary facts, multi-source answers, and explicit no-answer queries. The harness reports Recall@K, document Precision@K, MRR, binary nDCG@K, no-answer accuracy, citation/source-anchor coverage, and measured p50/p95/max query latency. Latency is diagnostic in the deterministic suite; hardware-specific 500 ms acceptance remains a benchmark gate rather than a timing-sensitive unit assertion. The deterministic suite has an opt-in real-model counterpart, Live/RetrievalEvalLiveTests, which scores the same metrics with a real embedder and real rerankers on llama-server and reports a config as INVALID instead of a number when its forced rerank degraded; run it with scripts/run-retrieval-eval-local.sh.
That invariant is a statement about the search response, not about what the user sees. Because storage_path is a GUID filename, a chunk with no heading trail would otherwise be labelled with a raw GUID — unreadable, and indistinguishable between two documents. The UI resolves the display name client-side: each hit carries its document_id, and KnowledgeSearchPanel joins that against the already-loaded documents list (which decrypts names for the same operator-authenticated viewer) to render 12-security-and-privacy.md › Encryption at rest, falling back to the API's title when the document is not in the loaded list. So do not "fix" a GUID-titled result by decrypting original_file_name in KnowledgeSearchService — the readable label already exists one layer up, and moving it into the response would also hand file names to a cloud model through the gated search_knowledge_base tool.
Without an embedding model installed the knowledge base cannot index anything — ingestion fails in KnowledgeChunkEmbedder with a content-free "not available" reason — so RecommendedEmbeddingModel gives a fresh node a one-click way out, exactly as RecommendedRerankerModel does for the strictly optional reranker. Both resolve through the same operator Hugging Face download path as any other GGUF (IGgufModelStore.EnsureModelAsync), and both model names carry a kind fragment (embed, reranker) so ModelKindDetector classifies them and keeps them out of the chat picker.
Embedding — nomic-ai/nomic-embed-text-v1.5-GGUF at F16, ~274 MB.
- That exact repository, not merely "some embedding model".
KnowledgeEmbeddingVectorPolicyapplies the versioned Nomic Matryoshka-512 transform only when the resolved model name containsnomic-embed-text-v1.5; every other model stays at its native width. Recommending the v1.5 repo therefore lands the node on the vector policy the shipped defaults (EmbeddingVectorMode = Matryoshka512) were designed around. A v2, or a different publisher, silently falls through to native width — still functional, but not what those defaults assume. F16rather than a small quant. The model is ~137M parameters, so full precision is only ~274 MB: the quantization saving is a few hundred megabytes at most, while embedding quality degrades in a way that is invisible at query time — retrieval silently returns worse chunks instead of failing. For a retrieval backbone every KB search depends on, the full-precision file is the right default. The pinnednomic-ai/nomic-embed-text-v1.5-GGUF:F16is live-verified: it downloads, spawns an embeddingllama-server, indexes, and retrieves correctly.- The download request leaves
GgufModelRequest.RoleatGgufRole.Embedding, so the supervisor spawns it with the embedding-role flags (--embeddings --pooling mean) rather than the chat ones. This is the one place the embedding recommendation legitimately differs from the reranker, which has no role of its own. RecommendedEmbeddingModel.ResolveExistingAsyncreads the local registry only in two steps, and the order is the point: the recommended repo wins when present, so the reported identity is stable; failing that, any installed embedding-named model counts, because that is exactly whatEmbeddingModelResolverwould pick, so downloading a second embedder would change nothing but the bandwidth bill. The fallback is ordered by name, so a node with several embedders reports the same one on every call rather than whatever the registry listed first.
Reranker — gpustack/bge-reranker-v2-m3-GGUF at Q4_K_M, ~440 MB, the quality/size sweet spot for a ~560M-parameter cross-encoder. There is no reranker GgufRole — the process split only distinguishes chat and embedding — so the request leaves GgufModelRequest.Role at GgufRole.Unknown and the reranker is identified by name, not by a stored role hint. RecommendedRerankerModel.ResolveExistingAsync is deliberately narrower than the embedding one: reranking is optional and choosing a reranker is an explicit operator act, so only this repo counts as "already installed" — another installed reranker is not one the operator asked for here.
Memory accounting. A cold embedder or reranker spawn passes the same capacity admission as chat launches (ILlamaServerPooledLaunchAdmission, implemented over CapacityService). The reranker is refused when memory does not allow it, and search stays on fusion order. The embedder is booked in the ledger but never refused on budget, because ingestion depends on it. Chat model-fit recommendations reserve the footprint of the configured reranker and llama.cpp embedder, against a freshly probed profile; a resident companion is skipped only when the budget is measured free VRAM, which already nets it out, so a recommended chat model leaves room for them.
The knowledge base is exposed to the agent loop as tools:
SearchKnowledgeBaseToolHandler— the agent issues a collection-scoped query and receives fused, hydrated hits (the sameIKnowledgeSearchServicepath), including thecollectionIdneeded for follow-up reads.ReadDocumentToolHandler— the agent reads one bounded document only when both its document id and collection id match.ReadSurroundingChunksToolHandler— the agent expands a specific hit with its neighboring chunks only inside the hit's collection.
These make the KB a retrieval-augmented generation (RAG) source the agent can consult mid-conversation. See Agent Mode for the tool registry.
KnowledgeBaseHub (KnowledgeIndexingNotifier) is a server-push-only hub: clients receive sanitized document status-change events; there are no client-callable server methods. It is protected with the same operator policy as the other local hubs because the indexing stream reveals which documents exist and are being processed. The React feature invalidates its document-list query on each event (notification-only). See API & Hubs.
Routes under knowledge/*, one endpoint class per file in Endpoints/Knowledge/V1/:
| Endpoint | Role |
|---|---|
UploadKnowledgeDocumentEndpoint |
Upload a document; returns immediately, ingestion runs async. |
ListKnowledgeDocumentsEndpoint |
List documents with their ingestion status. |
GetKnowledgeDocumentEndpoint |
One document's detail. |
DeleteKnowledgeDocumentEndpoint |
Delete a document and its chunks/index rows, then its encrypted blob. |
SearchKnowledgeEndpoint |
Hybrid search over the corpus (or a single document). |
ImportKnowledgeRepositoryEndpoint |
Import supported files from a registered local Git repository into one collection. |
ReindexKnowledgeDocumentEndpoint |
Re-run ingestion for one document. |
ReindexCorpusEndpoint |
Re-run ingestion for the whole corpus (e.g. after an embedding-model change). |
DownloadRecommendedRerankerEndpoint |
POST: one-click download of the recommended cross-encoder reranker via the same GGUF download coordinator operator HF downloads use; idempotent no-op if already installed or in flight. |
All endpoints are loopback/local-only, operator-authenticated, and secret-redacted — see Security & Privacy. They are surfaced to React via OpenAPI → hey-api; see API & Hubs.
Delete leaves nothing behind. KnowledgeDocumentPurgeService deletes every dependent row child-to-parent in one transaction (the node enforces the declared cascades, but chunk → section is SET NULL, and the purge reads each row's file location and returns what it removed — the FTS index is not a reason, since knowledge_document_chunks_ad fires for a cascade-removed chunk too) and removes the encrypted blob only after that commit — a failure there can therefore never leave a live row without its content. Because the rows are already gone, a failed blob delete is logged and still reported as a successful delete rather than a 500, and the file is reclaimed by KnowledgeBlobOrphanSweeper, a one-shot startup sweep that deletes any blob whose knowledge_documents row no longer exists. A repeat purge cannot do that job: it keys on the row it already deleted, returns 404, and never touches the disk.
src/features/knowledge/ (pages/, components/, hooks/, queries/, models/) renders the active-collection selector, collection-bound upload/document/search surfaces, source/provenance details, and the Development Mode registered-repository importer. It follows the standard client conventions (TanStack Query for server state, a SignalR hub that invalidates queries on push). See React Client.
- No document/chunk/query text is ever logged — failure logging is exception-type-only;
failure_reasonis a fixed content-free string. - Vectors compare only within one immutable model revision and dimension. A dimension mismatch is a
Faileddocument, not a silent wrong answer. Changing or replacing the embedding model requires a re-index (ReindexCorpusEndpoint). - Search degrades, never 500s, when the embedding model or reranker is unavailable.
- Results never expose the encrypted original file name —
Title/Sectioncome from the non-sensitive heading/storage path only. - The
Indexedtransition is atomic — a document is never half-visible to retrieval. - Every retrieval arm and hydration query applies the same collection scope — document filters never override namespace isolation.
- Embedding reuse is exact or disabled — content, parser/chunker versions, canonical vector identity, and dimension must all match.
- Local Runtime & Providers — the embedding/reranker models resolve through the same provider seams.
- Agent Mode — the KB search/read tools in the agent tool registry.
- Chat — file upload → chat attachments (a separate path from the KB corpus).
- Data & Persistence —
knowledge_documents, chunk/FTS/vector tables, encrypted blob store, migration. - API & Hubs —
/api/local/v1mapping, the indexing hub, OpenAPI → hey-api. - React Client — TanStack Query + SignalR conventions used by this feature.
- Security & Privacy — local-only endpoints, secret redaction, node-local privacy.
- Architecture Overview · Project Layout