Skip to content

Latest commit

 

History

History
165 lines (117 loc) · 36.5 KB

File metadata and controls

165 lines (117 loc) · 36.5 KB

Knowledge Base / RAG

Reviewed: 2026-09-15 · Code-grounded.

The Knowledge Base is a fully offline document store with hybrid retrieval. An operator uploads documents; the node extracts their text, chunks them on header boundaries, embeds each chunk with a local embedding model, and indexes everything into local SQLite with selective encryption — source document blobs and display names are encrypted at rest, while the extracted chunk text and its FTS search index are stored unencrypted locally. Retrieval fuses a lexical arm (SQLite FTS5 / BM25) and a semantic arm (vector cosine similarity) with Reciprocal Rank Fusion, optionally rescoring with a local cross-encoder reranker, and exposes the result to agents as a tool.

Provider-locality invariant (enforced default). Ingestion and embedding stay on-device by default (see KnowledgeBase:EmbeddingProviderName), and — as of the MED-004 hardening — the read-only knowledge tools (search_knowledge_base, read_document, read_surrounding_chunks) are offered only to node-local models (llama.cpp / Ollama). A cloud-hosted model (Codex OAuth, Azure Foundry) is not offered these tools, so document text, chunks, and the query are never handed to a third-party provider through a tool call. An operator can waive this by setting KnowledgeBase:AllowCloudModelAccess=true (default false) — an explicit, documented opt-in that acknowledges knowledge-base content may then reach the cloud provider.

The gate keys on the EFFECTIVE model — the one that actually runs the turn after any agent/profile pin is applied — not the turn's active model. So an agent, orchestration participant, or spawned sub-agent pinned to a cloud model is withheld the knowledge tools even when the conversation's active model is local (each effective model's locality is resolved through the shared IModelCapabilityResolver; participants and pins are classified individually, spawned children through the same resolver keyed on the child's pin). Retrieved chunk text — and its attacker-controlled metadata (title, section, source) — is additionally returned to the model fenced as untrusted data: a contentTrust: "untrusted-document" flag plus per-response nonce-delimited begin/end markers (a random marker suffix the document body cannot forge), so an instruction embedded in a document, title, or heading cannot read as a system directive or break out of the fence. Nothing reaches a log.

Attachments and workspace file tools are gated by the same effective-model locality, with a visible notice. Conversation attachments (inlined in plain chat, or staged into the AgentHome workspace and read via the coder list_files/read_file/search_text tools) are node-local private data. When the effective model that runs a turn is cloud-hosted — including when an agent/profile pin substitutes a cloud model even though the user's active pick is local — and the operator has not opted in via KnowledgeBase:AllowCloudModelAccess, the attachments are withheld: not staged for the file tools, not inlined into the prompt, and the coder file tools are withheld from the offer. The user gets a visible turn notice naming the effective model. (The earlier "user chose the conversation model" rationale does not hold once a pin substitutes the model, so attachments are gated exactly like the Knowledge Base rather than waved through.) The opt-in restores them. Whatever content does reach the model is fenced as untrusted data (file name inside the fence) so it is never read as instructions; the attachment fence nonce is derived from a server secret (HKDF over the node key), not the client-visible conversation id, so a client cannot forge the fence's closing marker.

Config key. KnowledgeBase:AllowCloudModelAccess (default false) is the single opt-in governing all node-local private-data exposure to a cloud model: the knowledge tools, the coder workspace file tools, and conversation attachments. It is named under KnowledgeBase for continuity; set it to true to allow a cloud model to receive any of that node-local content.

Where the code lives

Concern Project / path
Ingestion pipeline driver XE-Local-AI-Engine.Client.Application/Services/Knowledge/Implementation/KnowledgeIngestionService.cs (IKnowledgeIngestionService)
Background ingestion worker + dispatcher …/Services/Knowledge/Implementation/KnowledgeIngestionWorker.cs, KnowledgeIngestionDispatcher.cs
Header-boundary chunker …/Services/Knowledge/Implementation/HeaderBoundaryChunkingService.cs (IChunkingService)
Chunk embedder …/Services/Knowledge/Implementation/KnowledgeChunkEmbedder.cs (IKnowledgeChunkEmbedder)
Exact chunk-embedding reuse …/Services/Knowledge/Implementation/KnowledgeChunkEmbeddingCache.cs, KnowledgeChunkEmbeddingReuseStore.cs
Embedding model resolver + prefixer …/Services/Knowledge/Implementation/EmbeddingModelResolver.cs, KnowledgeEmbeddingPrefixer.cs
Hybrid search orchestrator …/Services/Knowledge/Implementation/KnowledgeSearchService.cs (IKnowledgeSearchService)
Optional-stage policy …/Services/Knowledge/AdaptiveRetrievalPolicy.cs
Lexical arm (FTS5) …/Services/Knowledge/Implementation/FtsSearch.cs (IFtsSearch)
Semantic arm (cosine) + factory …/Services/Knowledge/Implementation/ManagedCosineVectorSearch.cs, VectorSearchFactory.cs (IVectorSearch)
Encrypted document blob store …/Services/Knowledge/Implementation/KnowledgeDocumentBlobStore.cs (IKnowledgeDocumentBlobStore)
Document text extraction …/Services/DocumentIngestion/Implementation/DocumentTextExtractor.cs + Extraction/
Registered-repository import …/Services/Knowledge/Implementation/KnowledgeRepositoryImportService.cs
Scheduled stale-model reindex …/Services/Knowledge/Implementation/KnowledgeScheduledModelReindexWorker.cs
Agent-facing tools …/Services/Knowledge/Tools/Implementation/SearchKnowledgeBaseToolHandler.cs, ReadSurroundingChunksToolHandler.cs
SignalR notifier + hub XE-Local-AI-Engine.Client/Hubs/KnowledgeIndexingNotifier.cs, KnowledgeBaseHub.cs
Local endpoints XE-Local-AI-Engine.Client/Endpoints/Knowledge/V1/
React feature XE-Local-AI-Engine.Client.React/src/features/knowledge/
Retrieval evaluation corpus/harness XE-Local-AI-Engine.Client.Persistence.Tests/Knowledge/RetrievalEval/

These types hold a NodeChatDbContext and speak raw ADO from the application layer. That is a reasoned permanent exception to the store rule in Code conventions, not migration debt, and Architecture/DbContextUserAllowlist.txt records it per type. Three measured grounds: KnowledgeDocumentStatus is bound as a SQL parameter by half of them and is a shipped OpenAPI schema id and a SignalR parameter, so relocating the SQL would rename a generated client symbol; KnowledgeDocumentCatalogService.ResetStaleDocumentsToPendingAsync evaluates the embedding-model staleness policy between its SELECT and its UPDATEs, and KnowledgeDocumentBlobStore.UpdateRepositoryDocumentAsync holds its transaction open across an encrypted blob write and atomic rename, rolling database and file back together; and four of these interfaces are injected by host endpoints, which EndpointDependencyTests forbids for a persistence type. The two startup maintenance jobs (KnowledgeVectorNormalizationBackfillService, KnowledgeBlobOrphanSweeper) are listed with the other database-maintenance jobs instead.

Ingestion pipeline

KnowledgeIngestionService drives one document through the pipeline, advancing knowledge_documents.status at each transition and setting a content-free failure_reason on any step failure. Uploads are processed asynchronously by KnowledgeIngestionWorker so the upload endpoint returns as soon as the document is admitted to the queue and the operator watches status over the hub.

Upload ──▶ Uploaded ──▶ Extracting ──▶ Chunking ──▶ Embedding ──▶ Indexed
                 │            │             │             │            │
      encrypted blob     text extract   header      local model    atomic index
      (blob store)       (.pdf/.docx/    boundaries  batches        write (FTS + vectors)
                          plaintext)     chunker
  • Extraction — DocumentTextExtractor dispatches by extension: .pdf and .docx have dedicated readers (.docx via the pure-managed Open XML SDK); .html/.htm use the deterministic HtmlDocumentReader; Markdown, plaintext, structured text, logs, and supported source-code extensions use PlaintextDocumentReader. Markdown headings and fenced code blocks, HTML headings/paragraphs/lists/tables/code blocks, and code whitespace/symbol text are retained as ordered structural elements instead of being flattened prematurely. Extraction is bounded by DocumentExtractionLimits (a container format such as a .docx zip or .pdf can expand well beyond its on-disk size, so a decompression-bomb guard caps the extracted size).
  • Chunking — HeaderBoundaryChunkingService is an offline, deterministic chunker with no tokenizer or external package. It walks the document's ordered element stream: each header opens a new section (maintaining an H1 > H2 heading trail via a level stack); paragraphs, tables, and code blocks accumulate into their section body, which is split into token- and character-bounded, overlapping chunks (MaxChunkTokens, MaxChunkChars, ChunkOverlapChars). Heading trails are prepended to the embedding input as deterministic document context. Each chunk preserves offsets plus content kind, language, symbol, page, and source-path provenance where the parser can supply them. The same document and versioned parser/chunker settings yield the same sections, chunks, hashes, and identifiers. Sizing is token-aware: a section body is cut at whichever bound is reached first — the per-chunk token budget (MaxChunkTokens, tightened when a resolved embedding context window is known to that window minus a safety reserve covering the search_document: intent prefix and the model's own special tokens, and floored so a tiny or misconfigured window cannot collapse chunking to near-single-character windows) or the hard character ceiling (MaxChunkChars) — always breaking at the last whitespace before the limit so no word is cut. Token-awareness is what keeps a chunk plus its heading prefix inside the embedding model's context window and stops a byte-pair tokenizer splitting a max-sized chunk mid-token. The per-window token budget is reduced by the section's heading-trail cost, so the embedded heading-prefixed ContextualContent stays inside the window. Token counts come from ChunkTokenApproximation, a deterministic dependency-free approximation: weighted characters ÷ 4, with CJK and emoji weighted heavier so a token-dense script does not silently produce chunks several times the intended token size. The token budget can only tighten the effective size, never enlarge it past the character ceiling — plain ASCII prose therefore keeps the character ceiling as its binding bound and existing ASCII corpora need no reindex, while token-dense content and smaller embedder windows yield correspondingly smaller chunks.
  • Embedding — KnowledgeChunkEmbedder reuses the node-local embedding resolution path: it resolves the configured embedding provider/model (KnowledgeBaseOptions.EmbeddingProviderName / EmbeddingModelName) and generates in batches (MaxEmbeddingBatchSize). After provider generation, ingestion and query share KnowledgeEmbeddingVectorPolicy. A confidently resolved nomic-embed-text-v1.5 defaults to the versioned Matryoshka recipe (full-vector population layer norm with epsilon 1e-5 → first 512 components → L2 normalization); other models stay native. The provider's immutable installed-weight digest/revision is hashed into the resolved inventory fingerprint, so replacing weights under the same Ollama tag or GGUF name changes the canonical identity. The exact resolved model revision + transform/version + width is persisted and filtered during search, with dimension checked separately. When that identity and width are known before inference, ingestion reuses an exact vector keyed by contextual-content hash, parser version, chunker version, canonical vector identity, and dimension. A process-wide TTL/entry/byte-bounded RAM cache and in-flight coalescer first consult a short-lived scoped durable lookup of committed indexed rows; unknown-width providers retain the normal embed path rather than risk a false hit.
  • Indexing — the final Indexed transition is performed atomically by IKnowledgeIndexWriter so a document is never half-visible. All failure logging is exception-type-only — no chunk or document text reaches a log.
  • Admission control (backpressure) — the async queue is bounded (KnowledgeIngestionDispatcher: a single-reader channel, capacity 256; the worker further caps concurrency with a SemaphoreSlim). Admission is non-blocking: when the queue is full, the upload/reindex endpoint returns 503 with Retry-After: 5 instead of holding the request or growing the backlog, and admission is idempotent (a document already queued or in flight is not re-enqueued, so it is never processed twice concurrently). A 503 is a retryable busy signal, not data loss — the blob is already persisted, so a document whose ingestion a full queue previously rejected (persisted-but-unindexed) is re-enqueued on a later re-upload of the same file. Accept/reject counts and live queue depth are published on the XE.Node meter. The synchronous chat-attachment path has its own equivalent gate (DocumentExtractionAdmissionGate), also 503 + Retry-After when at capacity — see the busy-admission note in API & Hubs.
  • Worker lifecycle — the queue is an in-memory channel, so a document left mid-pipeline by a crash or hard stop would be stuck non-terminal forever: on start KnowledgeIngestionWorker resets every non-terminal document to Pending and re-dispatches it before draining new work. At shutdown it stops reading the queue and awaits the tracked in-flight documents within a bounded drain window (ShutdownDrainTimeoutSeconds) before disposal proceeds, so the scope factory and the semaphore are never disposed under a running document. Per-document work runs on a drain-deadline token cancelled only when that window elapses — ordinary operation never cancels a document mid-write, yet a hung document cannot block shutdown forever, and anything unfinished is left non-terminal and re-queued on the next start. A separate drain sweep re-admits documents still Pending once the admission queue empties, so uploads a full queue rejected are picked up as capacity frees; admission is idempotent, so nothing is ingested twice.
  • Scheduled index migration — KnowledgeScheduledModelReindexWorker periodically discovers documents whose stored canonical vector identity, immutable model revision, parser version, or chunker version no longer matches the active pipeline and admits them to the same bounded background queue. Deterministic parser/chunker migrations still mark rows stale during a transient model-provider outage; model identity comparisons require a confident provider resolution — an installed model actually matched, never the configured-name fallback that an unreachable provider or an unmatched configured name produces. On a llama.cpp node the stored embedding_model is a resolved GGUF name that never equals the plain configured name, so comparing against that fallback during an outage would flag, and reset, the entire indexed corpus instead of leaving it untouched. Only Indexed rows can be stale at all: a non-indexed row still carries the upload-time placeholder name, which is not a vector identity. It does not bypass normal queue pressure or atomic indexing. ScheduledModelReindexEnabled and ScheduledModelReindexIntervalMinutes control the scheduler. Bumping KnowledgeIndexVersions.Parser re-queues every indexed document and re-embeds all of its chunks, because the parser version is part of the embedding reuse key.

Matryoshka rollback procedure. Set KnowledgeBase:EmbeddingVectorMode to Native, invoke the normal corpus reindex endpoint, and wait until the catalog reports no stale documents. Only then deploy a binary that predates canonical vector identities. Switching the binary first is unsafe because old model-name-only search code cannot distinguish the 512-wide transformed rows from native vectors. Switching either direction changes the canonical identity immediately, so old projections stay excluded/stale until the normal per-document or full-corpus reindex rebuilds them.

Collections and repository ingestion

Every document, section, chunk, FTS row, vector candidate, and hydrated hit belongs to a canonical CollectionId. Missing collection input maps to DEFAULT for backward compatibility; explicit ids are trimmed, upper-cased, limited to 128 ASCII letters/digits plus -, _, and ., and applied independently to both lexical and dense retrieval. A DocumentId from another collection returns no results rather than weakening the collection boundary. This is a local namespace/isolation boundary, not a multi-tenant authorization system.

Development Mode can import a previously registered local Git repository through ImportKnowledgeRepositoryEndpoint. The request carries only the registered selected-folder id and collection id—never an arbitrary path. KnowledgeRepositoryImportService revalidates the registered folder as a local Git top-level, refuses reparse/symlink escapes, skips sensitive/generated directories and files, enforces per-file plus aggregate byte limits, revalidates the opened file before accepting its bytes, preserves only relative source paths, and sends eligible source files through the normal encrypted-blob/background-ingestion pipeline. Repository identity is collection + normalized source path: re-import updates a changed path in place, identical bytes at two paths keep distinct provenance, unchanged paths are not requeued, and a successfully completed scan purges repository documents whose paths were deleted or renamed. Uploads retain collection-scoped content-hash deduplication.

Hybrid retrieval

KnowledgeSearchService.SearchAsync is the retrieval heart. A KnowledgeSearchRequest carries the untrusted Query, a Limit, a normalized collection scope (defaulting to DEFAULT), an optional DocumentId scope, and an ExpandNeighbors flag. The flow:

  1. Embed the query with the current model, using the query-intent prefix (KnowledgeEmbeddingPrefixer).
  2. Two arms retrieve candidates in parallel:
    • Lexical — FtsSearch runs a BM25-ranked MATCH over weighted FTS5 fields for source path, heading path, symbol, and body content (the raw-SQL path; the query is escaped before it reaches FTS). Exact identifiers, file paths, and symbols therefore receive more weight than repeated body terms. Collection and optional document scopes are enforced in the same query through UNINDEXED identity columns.
    • Semantic — the collection- and model-scoped IVectorSearch (ManagedCosineVectorSearch) scores chunk embeddings by cosine similarity within the active canonical model identity/dimension. Each arm fetches a candidate pool (CandidatePoolMultiplier × limit, floored at MinimumCandidatePool) so fusion has enough overlap material. The two arms are launched together, but only the lexical arm touches the request-scoped SQLite connection — the query-embedding arm talks solely to the embedding provider process — so overlapping them can never run two commands on that non-thread-safe connection at once. Every DB-bound read that consumes the embedding (the vector scan, hydration, neighbor expansion) runs sequentially afterwards on that one connection; truly concurrent execution of both DB arms would need a second connection/scope and is deliberately not taken. The win is overlapping the embedding round trip, typically the dominant latency, with the FTS query. The vector arm is filtered by the same resolved model name the query was embedded with, so query vectors are only ever compared against chunk vectors built by that identical model — a later same-dimension model swap changes the name and excludes the now-incompatible old vectors instead of silently mis-comparing them.
  3. Fuse the two ranked lists with Reciprocal Rank Fusion (IRankingFusionService). The fused set is the union of every chunk id either arm returned. Two strategies exist. Rrf is the classic score-agnostic fusion that uses rank position only — the graceful fallback and the comparison baseline. ScoreAware min-max normalizes each arm's scores within that arm, so incomparable scales (BM25 magnitude vs cosine) never blend directly, and uses the normalized value to tilt the RRF contribution multiplicatively by up to a configured weight, so a rank whose arm score sits far above its arm's floor outranks an equally-ranked but marginal competitor. ScoreAware degrades to pure RRF for any arm whose scores are empty, single, constant, or non-finite, because normalization carries no signal there. Callers orient incomparable raw scores before fusion so that higher always means more relevant — FTS5 BM25, which is more-negative-for-stronger, is negated.
  4. Conditionally rerank the fused candidate pool with a local cross-encoder reranker (IRerankerClient, model KnowledgeBaseOptions.RerankerModelName) before the top-limit cut. With AdaptiveRerankingEnabled (default), the deterministic policy skips the optional model when both retrieval arms agree on the top chunk, the candidate set is too small, or at least 80% of RetrievalLatencyBudgetMilliseconds (default 500 ms) has elapsed; ambiguous candidates still rerank. The same budget, measured from search start, is the hard deadline for the rerank stage. Search never spawns the reranker: LlamaServerRerankerClient only uses a process that is already running, so a cold reranker returns null at once (degrade reason cold), the search returns fusion order, and KnowledgeSearchService asks IKnowledgeModelPrewarmer for a background warm; start-up time is therefore never inside the budget. Scoring is abandoned, not cancelled, at the deadline: llama-server keeps scoring a cancelled request anyway, so the HTTP call finishes in the background under the client's own 5–30 s timeout, and while it is outstanding every new rerank for that model falls back at once (degrade reason busy). Deadline expiry degrades to fusion order; caller cancellation still propagates. Disabling the adaptive gate restores always-attempt-rerank behavior for controlled benchmarks, but not an unbounded wait. Every hit carries a ScoreKind naming the scale of its Score, one kind per result because reranking is all-or-nothing: Fusion is the fused RRF score (roughly 0.01–0.06), Rerank the raw cross-encoder relevance, set only when reranking succeeds (unbounded and model-specific: bge roughly −11…+7, qwen3 0…1). Every fallback (reranker off, gate skip, null or count-mismatched response, deadline expiry) stays Fusion. Neither is a probability, and scores compare only within one result. The REST hit (scoreKind: "Fusion"/"Rerank"), the agent tool JSON and persisted chat sources carry it; a chat source saved before it existed loads with a null kind.
  5. Hydrate the selected chunks over the raw-SQL path and, when ExpandNeighbors is set, expand each hit with its surrounding neighbor chunks (NeighborWindow = 1 each side).

Pre-warm. KnowledgeModelPrewarmer (singleton) ensures the configured reranker and, on the llama.cpp provider, the resolved embedder through EnsureRunningAsync. InvocationRunner requests it once per chat or agent turn whose offer contains search_knowledge_base (with AgentToolsEnabled on), after the turn's own model is ready so the two launches never race; a search whose rerank came back empty, malformed or past the retrieval budget requests it again, and that warm's reuse path re-probes a live reranker's liveness, so one that accepts and hangs is torn down and respawned. Requests coalesce into one warm in flight; after a failed or refused warm, requests are ignored for 30 s so repeated searches do not repeat the capacity probe. A successful warm sets no cooldown: re-ensuring a running process is a cheap reuse that also refreshes its idle timestamp, since pooled processes are still reaped after the idle TTL.

Graceful degradation is a contract: if the embedding model or the reranker is unavailable, search degrades (lexical-only / fusion-order) rather than failing. Each hit carries collection, document, chunk, relative source path, heading/section, page, offsets, content kind, language, and symbol provenance where available. A hit's Title/Section are derived from the non-sensitive heading_path/storage_path, so a result never exposes the encrypted original file name.

Retrieval evaluation

Client.Persistence.Tests/Knowledge/RetrievalEval/ runs the real ingestion, SQLite FTS/vector, fusion, hydration, and optional reranker path against deterministic embeddings. Its representative corpus covers English and German semantic questions, exact code identifiers/paths, lexical distractors, long-document boundary facts, multi-source answers, and explicit no-answer queries. The harness reports Recall@K, document Precision@K, MRR, binary nDCG@K, no-answer accuracy, citation/source-anchor coverage, and measured p50/p95/max query latency. Latency is diagnostic in the deterministic suite; hardware-specific 500 ms acceptance remains a benchmark gate rather than a timing-sensitive unit assertion. The deterministic suite has an opt-in real-model counterpart, Live/RetrievalEvalLiveTests, which scores the same metrics with a real embedder and real rerankers on llama-server and reports a config as INVALID instead of a number when its forced rerank degraded; run it with scripts/run-retrieval-eval-local.sh.

That invariant is a statement about the search response, not about what the user sees. Because storage_path is a GUID filename, a chunk with no heading trail would otherwise be labelled with a raw GUID — unreadable, and indistinguishable between two documents. The UI resolves the display name client-side: each hit carries its document_id, and KnowledgeSearchPanel joins that against the already-loaded documents list (which decrypts names for the same operator-authenticated viewer) to render 12-security-and-privacy.md › Encryption at rest, falling back to the API's title when the document is not in the loaded list. So do not "fix" a GUID-titled result by decrypting original_file_name in KnowledgeSearchService — the readable label already exists one layer up, and moving it into the response would also hand file names to a cloud model through the gated search_knowledge_base tool.

Recommended embedding and reranker models

Without an embedding model installed the knowledge base cannot index anything — ingestion fails in KnowledgeChunkEmbedder with a content-free "not available" reason — so RecommendedEmbeddingModel gives a fresh node a one-click way out, exactly as RecommendedRerankerModel does for the strictly optional reranker. Both resolve through the same operator Hugging Face download path as any other GGUF (IGgufModelStore.EnsureModelAsync), and both model names carry a kind fragment (embed, reranker) so ModelKindDetector classifies them and keeps them out of the chat picker.

Embedding — nomic-ai/nomic-embed-text-v1.5-GGUF at F16, ~274 MB.

  • That exact repository, not merely "some embedding model". KnowledgeEmbeddingVectorPolicy applies the versioned Nomic Matryoshka-512 transform only when the resolved model name contains nomic-embed-text-v1.5; every other model stays at its native width. Recommending the v1.5 repo therefore lands the node on the vector policy the shipped defaults (EmbeddingVectorMode = Matryoshka512) were designed around. A v2, or a different publisher, silently falls through to native width — still functional, but not what those defaults assume.
  • F16 rather than a small quant. The model is ~137M parameters, so full precision is only ~274 MB: the quantization saving is a few hundred megabytes at most, while embedding quality degrades in a way that is invisible at query time — retrieval silently returns worse chunks instead of failing. For a retrieval backbone every KB search depends on, the full-precision file is the right default. The pinned nomic-ai/nomic-embed-text-v1.5-GGUF:F16 is live-verified: it downloads, spawns an embedding llama-server, indexes, and retrieves correctly.
  • The download request leaves GgufModelRequest.Role at GgufRole.Embedding, so the supervisor spawns it with the embedding-role flags (--embeddings --pooling mean) rather than the chat ones. This is the one place the embedding recommendation legitimately differs from the reranker, which has no role of its own.
  • RecommendedEmbeddingModel.ResolveExistingAsync reads the local registry only in two steps, and the order is the point: the recommended repo wins when present, so the reported identity is stable; failing that, any installed embedding-named model counts, because that is exactly what EmbeddingModelResolver would pick, so downloading a second embedder would change nothing but the bandwidth bill. The fallback is ordered by name, so a node with several embedders reports the same one on every call rather than whatever the registry listed first.

Reranker — gpustack/bge-reranker-v2-m3-GGUF at Q4_K_M, ~440 MB, the quality/size sweet spot for a ~560M-parameter cross-encoder. There is no reranker GgufRole — the process split only distinguishes chat and embedding — so the request leaves GgufModelRequest.Role at GgufRole.Unknown and the reranker is identified by name, not by a stored role hint. RecommendedRerankerModel.ResolveExistingAsync is deliberately narrower than the embedding one: reranking is optional and choosing a reranker is an explicit operator act, so only this repo counts as "already installed" — another installed reranker is not one the operator asked for here.

Memory accounting. A cold embedder or reranker spawn passes the same capacity admission as chat launches (ILlamaServerPooledLaunchAdmission, implemented over CapacityService). The reranker is refused when memory does not allow it, and search stays on fusion order. The embedder is booked in the ledger but never refused on budget, because ingestion depends on it. Chat model-fit recommendations reserve the footprint of the configured reranker and llama.cpp embedder, against a freshly probed profile; a resident companion is skipped only when the budget is measured free VRAM, which already nets it out, so a recommended chat model leaves room for them.

Agent tool surface

The knowledge base is exposed to the agent loop as tools:

  • SearchKnowledgeBaseToolHandler — the agent issues a collection-scoped query and receives fused, hydrated hits (the same IKnowledgeSearchService path), including the collectionId needed for follow-up reads.
  • ReadDocumentToolHandler — the agent reads one bounded document only when both its document id and collection id match.
  • ReadSurroundingChunksToolHandler — the agent expands a specific hit with its neighboring chunks only inside the hit's collection.

These make the KB a retrieval-augmented generation (RAG) source the agent can consult mid-conversation. See Agent Mode for the tool registry.

SignalR notifications

KnowledgeBaseHub (KnowledgeIndexingNotifier) is a server-push-only hub: clients receive sanitized document status-change events; there are no client-callable server methods. It is protected with the same operator policy as the other local hubs because the indexing stream reveals which documents exist and are being processed. The React feature invalidates its document-list query on each event (notification-only). See API & Hubs.

Endpoints

Routes under knowledge/*, one endpoint class per file in Endpoints/Knowledge/V1/:

Endpoint Role
UploadKnowledgeDocumentEndpoint Upload a document; returns immediately, ingestion runs async.
ListKnowledgeDocumentsEndpoint List documents with their ingestion status.
GetKnowledgeDocumentEndpoint One document's detail.
DeleteKnowledgeDocumentEndpoint Delete a document and its chunks/index rows, then its encrypted blob.
SearchKnowledgeEndpoint Hybrid search over the corpus (or a single document).
ImportKnowledgeRepositoryEndpoint Import supported files from a registered local Git repository into one collection.
ReindexKnowledgeDocumentEndpoint Re-run ingestion for one document.
ReindexCorpusEndpoint Re-run ingestion for the whole corpus (e.g. after an embedding-model change).
DownloadRecommendedRerankerEndpoint POST: one-click download of the recommended cross-encoder reranker via the same GGUF download coordinator operator HF downloads use; idempotent no-op if already installed or in flight.

All endpoints are loopback/local-only, operator-authenticated, and secret-redacted — see Security & Privacy. They are surfaced to React via OpenAPI → hey-api; see API & Hubs.

Delete leaves nothing behind. KnowledgeDocumentPurgeService deletes every dependent row child-to-parent in one transaction (the node enforces the declared cascades, but chunk → section is SET NULL, and the purge reads each row's file location and returns what it removed — the FTS index is not a reason, since knowledge_document_chunks_ad fires for a cascade-removed chunk too) and removes the encrypted blob only after that commit — a failure there can therefore never leave a live row without its content. Because the rows are already gone, a failed blob delete is logged and still reported as a successful delete rather than a 500, and the file is reclaimed by KnowledgeBlobOrphanSweeper, a one-shot startup sweep that deletes any blob whose knowledge_documents row no longer exists. A repeat purge cannot do that job: it keys on the row it already deleted, returns 404, and never touches the disk.

React feature

src/features/knowledge/ (pages/, components/, hooks/, queries/, models/) renders the active-collection selector, collection-bound upload/document/search surfaces, source/provenance details, and the Development Mode registered-repository importer. It follows the standard client conventions (TanStack Query for server state, a SignalR hub that invalidates queries on push). See React Client.

Invariants a maintainer must respect

  1. No document/chunk/query text is ever logged — failure logging is exception-type-only; failure_reason is a fixed content-free string.
  2. Vectors compare only within one immutable model revision and dimension. A dimension mismatch is a Failed document, not a silent wrong answer. Changing or replacing the embedding model requires a re-index (ReindexCorpusEndpoint).
  3. Search degrades, never 500s, when the embedding model or reranker is unavailable.
  4. Results never expose the encrypted original file name — Title/Section come from the non-sensitive heading/storage path only.
  5. The Indexed transition is atomic — a document is never half-visible to retrieval.
  6. Every retrieval arm and hydration query applies the same collection scope — document filters never override namespace isolation.
  7. Embedding reuse is exact or disabled — content, parser/chunker versions, canonical vector identity, and dimension must all match.

Related pages