This is a Docker Compose-based AI stack managed via systemd. The stack provides local LLM inference (ollama), cloud API routing (LiteLLM), unified routing/load balancing (Olla), and Obsidian vault RAG (retriever). The primary AI interface is OpenCode (CLI + Obsidian sidebar plugin).
NEVER run docker compose up directly — it bypasses GPU overlay selection, VaultWarden
placeholder resolution, and Olla config generation. The ollama container will start without
GPU access (CPU-only mode) if the NVIDIA/Arc overlay is not applied.
# ✅ Correct — always use one of these:
sudo systemctl restart ai-stack.service # full stack via systemd
bash ./start.sh -d # full stack detached
bash ./start.sh -d --no-deps <service> # single service rebuild
# ❌ Never — bypasses GPU overlay and placeholders:
docker compose up -d # WRONG
docker compose up -d --no-deps ollama # WRONG — ollama loses GPU
docker compose build router && docker compose up -d router # WRONG
Root cause: start.sh auto-detects GPU type (NVIDIA/Arc/CPU) and passes
-f docker-compose.yml -f docker-compose.nvidia.yml (or .arc.yml) to compose.
Raw docker compose only loads docker-compose.yml (CPU-only for ollama).
GPU config on cluster-llm: RTX 3090 Ti (24 GB VRAM). GPU_TYPE=nvidia is pinned
in .env so detection never fails, but the overlay must still be passed via start.sh.
# Start / stop / restart (via systemd, preferred)
sudo systemctl start|stop|restart ai-stack.service
# Direct compose — MUST go through start.sh (see warning above)
./start.sh # foreground
./start.sh -d # detached
./start.sh -d --no-deps router # rebuild single service
./start.sh down # tear down
# Regenerate Olla config after changing .env OLLAMA_REMOTE_* entries
./scripts/generate-olla-config.sh
# Resolve Bitwarden/VaultWarden placeholders in .env (auto-runs in start.sh)
./scripts/resolve-vaultwarden.sh # resolve in-place
./scripts/resolve-vaultwarden.sh --dry-run # preview only
# Discover other Ollama nodes on the LAN
./scripts/discover-herd.sh # prompt before writing
./scripts/discover-herd.sh --apply # write without prompt
./scripts/discover-herd.sh --dry-run # scan only
# Sync models across cluster nodes (LAN speed if SSH keys deployed)
./scripts/sync-models.sh # sync all nodes
./scripts/sync-models.sh 10.10.0.212 # sync one node
./scripts/sync-models.sh --dry-run # preview only
./scripts/sync-models.sh --ssh-key ~/.ssh/id_ed25519 # with SSH key
# Herd observability — Shepherd (per-node sidecar + control-plane dashboard)
# Shepherd supersedes the earlier Apostle agent — see shepherd/README.md
scripts/shepherd-auto-deploy.sh node # deploy shepherd-node on this peer
scripts/shepherd-auto-deploy.sh control # deploy shepherd-control (one node only)
scripts/shepherd-auto-deploy.sh both # both (cluster-llm pattern)
# Endpoints once running:
# GET http://localhost:40116/herd/metrics — per-node snapshot
# GET http://localhost:40117/ — central dashboard (D3 + SSE)
# Discover AI services across all networks (LAN + VPN)
./scripts/discover-network.sh # interactive
./scripts/discover-network.sh 10.10.0.201:11434 # seed(s) as args, prompts
./scripts/discover-network.sh --apply # add all discovered
./scripts/discover-network.sh --dry-run # scan only
# Check retriever status
curl localhost:42000/health
# Search the vault
curl -X POST localhost:42000/search -H 'Content-Type: application/json' \
-d '{"query":"what did I write about networking?"}'
# Force vault reindex
curl -X POST localhost:42000/reindex
# CI validation (run locally before push)
docker compose config --quiet
shellcheck -s bash scripts/*.sh install.sh
# GPU pre-flight check
./scripts/check-arc-gpu.shAll traffic flows through Olla (port 40114) as the unified LLM router:
OpenCode (CLI + Obsidian plugin)
├── tool: retriever :42000 → sqlite-vec + FTS5 hybrid search over vault
├── provider: Olla :40115 → Smart Router (auto-selects local model)
│ → ollama :11434 (local LLM)
│ → OLLAMA_REMOTE_* nodes (LAN, optional)
└── provider: LiteLLM :4000 → Claude (Anthropic), Gemini (Google)
| Directory/File | Purpose |
|---|---|
docker-compose.yml |
Core stack: ollama, litellm, olla, router, retriever |
install.sh |
Preflight → create volumes → install systemd → start stack → pull models (prompts for Bitwarden setup) |
retriever/ |
Obsidian vault RAG: FastAPI + sqlite-vec + watchdog. Hybrid search via FTS5 + vector embeddings. |
scripts/generate-olla-config.sh |
Reads OLLAMA_REMOTE_* from .env → writes proxy/olla.yaml |
scripts/discover-herd.sh |
mDNS + subnet scan for other Ollama nodes on LAN |
scripts/check-arc-gpu.sh |
GPU pre-flight: detects card0/card1 drift, updates .env, used as ExecStartPre |
scripts/resolve-vaultwarden.sh |
Resolves <vaultwarden:path> placeholders in .env via bw CLI |
scripts/apostle.py |
(legacy — superseded by Shepherd) Earlier self-aware model agent; retained for reference until removed |
scripts/models.yaml |
Model catalog with 16 entries, RAM/disk/tools/priority/role metadata |
shepherd/ |
Herd observability service — per-node sidecar + control-plane dashboard. See shepherd/README.md |
scripts/shepherd-auto-deploy.sh |
One-shot or cron-driven deploy of shepherd-node / shepherd-control on a peer |
PLANS.md |
Original Project Apostle plan + status block flagging supersession by current architecture (Olla + Router + LiteLLM + Shepherd) |
router/ |
Smart Model Router: content-based model selection between OpenCode and Olla |
proxy/litellm_config.yaml |
Static LiteLLM model registry (Claude, Gemini models) |
.envis the single source of truth for LAN node addresses, GPU paths, API keys, and ports. Scripts parseOLLAMA_REMOTE_*vars — do not hardcode IPs in compose or scripts.proxy/olla.yamlis auto-generated — never edit directly. Regenerate viascripts/generate-olla-config.sh. It's in.gitignore.- GPU card node drifts (
card0vscard1) on Meteor Lake reboots.renderD128is stable.check-arc-gpu.shdetects and corrects via systemdExecStartPre. - Retriever volume is managed by compose (
retriever-data). Contains the sqlite-vec database with embeddings and FTS5 index. OLLAMA_KEEP_ALIVE=-1keeps models resident in shared system RAM (Intel iGPU uses shared memory, not dedicated VRAM).- VaultWarden integration is optional —
install.shprompts to set it up. If declined,.envstores API keys in plaintext (standard practice for local-only stacks). - Never commit secrets —
.envis in.gitignore. The installer generatesLITELLM_MASTER_KEYand stores it in Bitwarden automatically. Anthropic/Gemini keys are stored as<vaultwarden:...>placeholders and resolved at runtime.
The retriever service replaces Khoj + PostgreSQL with a lightweight, API-only service:
- Vector store: sqlite-vec (embedded SQLite extension, file-based, no separate DB)
- Keyword search: SQLite FTS5 (BM25 scoring)
- Hybrid search: Reciprocal Rank Fusion (RRF) combining vector + keyword results
- Embeddings:
nomic-embed-textvia Olla → ollama - Indexing: Full scan on startup, then watchdog (inotify) for live changes
- API:
POST /search,POST /reindex,GET /health
Configuration via .env:
RETRIEVER_PORT=42000
RETRIEVER_VAULT_PATH=/home/user/obsidian
RETRIEVER_EMBED_MODEL=nomic-embed-text
RETRIEVER_CHUNK_SIZE=512
RETRIEVER_CHUNK_OVERLAP=64
pytest.inisetstestpaths = testsandasyncio_mode = auto- Python test deps:
requirements-dev.txt(pytest, pytest-asyncio, httpx, pydantic) - CI runs:
docker compose config --quiet, shellcheck,.env.examplecredential scan
- CI (
.github/workflows/ci.yml): validates compose syntax, shellchecks scripts, scans.env.examplefor leaked credentials - Release (
.github/workflows/release.yml): triggered onv*.*.*tags, extracts notes fromCHANGELOG.md - Branch:
main
The project includes two custom tools for Obsidian vault search:
vault-search— search the entire vault for notes matching a queryvault-search_per_source— search within a specific file or subdirectory
Both tools call the retriever service (:42000) and return file paths, content snippets, and scores. Use them when asked to find information in notes.
- LiteLLM healthcheck URL is
/health/liveness(not "liveliness") install.shmodel pull is interactive (prompts y/N) — non-headless- Retriever depends on Olla for embeddings — ensure Olla is healthy before retriever starts
- Vault is mounted read-only (
:ro) — retriever never modifies notes discover-herd.shrequiresavahi-daemonrunning on the host for mDNS discoverybw login --apikeyis incompatible with self-hosted VaultWarden (nouserDecryptionOptionsin response). Use interactivebw loginor an existing unlocked session forresolve-vaultwarden.sh.install.shauto-generates theLITELLM_MASTER_KEY, creates alitellm-master-keyitem in your Bitwarden vault, and writes a<vaultwarden:...>placeholder to.env. Foranthropic-api-keyandgemini-api-key, you must create those items manually via the Bitwarden web vault.