Vera is a local-first AI home assistant with voice, wake word detection, MCP tools, OpenClaw integration, and a polished browser UI. It is designed to run on a Raspberry Pi or Mini-PC and speaks with streaming, low-latency TTS.
| Piece | Tech |
|---|---|
| Core server | Node.js + TypeScript (ESM), Express, ws |
| LLM | Gemini gemini-3.5-flash-lite (primary) → OrcaRouter → OpenRouter failover chain |
| Voice in | Gemini Live streaming transcription (gemini-3.5-transcribe-live) |
| Voice out | Fish Audio S2.1 → Fish S2.1 free via OpenRouter → Gemini TTS (auto failover) |
| Tools | MCP servers (stdio + HTTP) via @modelcontextprotocol/sdk |
| Automation | OpenClaw gateway (spawned by Vera) + headless task delegation |
| Wake word | Headless Python worker, local microWakeWord TFLite model (trained "Hey Vera" model included) |
| UI | Vanilla JS/HTML/CSS — chat, terminal, avatar, system monitor, MCP manager (no build step) |
| Desktop | Optional Electron wrapper (npm run desktop) or Chromium kiosk (scripts/kiosk.sh) |
# 1. Install (Node >= 22.22.3)
npm install
# 2. Configure
cp .env.example .env
# edit .env — see "API keys" below for what is actually required
# 3. Run
npm start # build + serve on http://localhost:3000
npm run dev # or: auto-reloading dev server (tsx watch)
npm run typecheck # typecheck without emittingOn boot Vera starts the OpenClaw gateway, the MCP registry servers, the system
monitor, and (unless disabled) the wake word worker. Open the UI at
http://localhost:3000 — the greeting is spoken once the browser sends its
first message.
Everything runs locally except the cloud AI calls you configure. Keys live in
.env (never commit it):
| Key | Needed for |
|---|---|
ORCAROUTER_API_KEY or OPENROUTER_API_KEY |
Text chat (at least one — these enable the LLM) |
GEMINI_API_KEY |
Primary chat model, all speech recognition (STT), and fallback TTS |
FISH_API_KEY |
Primary TTS voice (Fish Audio S2.1); free tier works |
Text chat works with just an OrcaRouter/OpenRouter key. Voice needs Gemini for
speech-to-text; the speaking voice then fails over Fish → OpenRouter → Gemini
automatically, so a single GEMINI_API_KEY plus one LLM key is a working setup.
.env.example documents every variable. The ones you'll most likely touch:
PORT— UI/server port (default 3000).FISH_REFERENCE_ID/FISH_ICARUS_REFERENCE_ID— custom cloned voices for Vera / Icarus.WAKE_ENABLED,WAKE_MODE,WAKE_WORD,WAKE_MODEL,WAKE_THRESHOLD— wake word (below).OPENCLAW_MODE—spawn(Vera launches the gateway, default) orexternal(attach to a gateway you started yourself withopenclaw gateway).
Glass theme: pure black background, translucent panels with glowing white outlines, gold accents (user messages, speaking orb, live mic).
- Main screen — the talking orb, chat history, system monitor strip, and a floating translucent dock at the bottom. No top bar, no tabs.
- Dock — persona toggle (Vera ⇄ Icarus), voice output toggle, mic button, chat view, terminal overlay, system monitor toggle, MCP servers overlay, send.
- Terminal — mirrors every server log line live (recent history is replayed to new clients).
- MCP — list / add / remove MCP servers at runtime.
Text chat works with or without voice. The mic button opens the browser mic;
Vera speaks replies through a streaming TTS chain, and every spoken frame is
counted by the voice: counter in the bottom-right of the UI.
Two assistants ship by default, each with her/his own prompt, voice, theme, and
conversation context: Vera (warm, witty) and Icarus (dry, precise).
Switch by saying "Icarus mode" / "Vera mode" / "Icarus off", or with the dock
toggle. A switch resets context, relabels the chat, and is announced with a
spoken greeting. Give Icarus his own voice with FISH_ICARUS_REFERENCE_ID and
GEMINI_ICARUS_TTS_VOICE.
Every message is routed into a lane that shapes the reply: Quick, Smart Home, Coding, Business, Research. A regex heuristic routes most messages instantly; ambiguous ones go through an LLM router prompt. The Quick lane runs with no tools attached (tool definitions measurably slow fast models' first token from ~0.6s to ~13s); other lanes can call MCP tools and OpenClaw delegation, for up to 3 tool rounds per turn. When the model goes straight to work, you get a spoken "on it" acknowledgement first, then the answer. Conversation history is bounded to the last 40 turns per connection.
Requests walk this chain in order; the first provider that streams wins:
gemini/gemini-3.5-flash-lite— primary (needsGEMINI_API_KEY, fast, free tier)orcarouter/deepseek/deepseek-v4-flash-freeorcarouter/free- Steps 2–3 again after a 4 s delay (usage-limit backoff)
openrouter/minimax/minimax-m2.7(cheap paid — the old:freeslug was removed upstream)openrouter/free(OpenRouter's auto free router)
If all steps fail, the error is logged and shown in chat.
Failover is aggressive and visible in the terminal panel:
- Rate-limited providers are skipped for up to 30 minutes (honoring the
provider's own
retry-afterhint when it sends one) instead of being retried mid-conversation. - Model slugs that 404 or are "unavailable for free" are skipped for an hour.
- A provider that accepts the request but never streams is cooled down for 60 s.
- Streams are protected by hang watchdogs: first chunk must arrive within 12 s (30 s when tools are attached), chunks must keep flowing (15 s stall limit), and the total stream is capped at 180 s.
- The OpenAI SDK's internal retries are disabled so failover happens here, visibly, in the log.
browser mic (PCM16 16 kHz) ──ws──▶ Gemini Live STT ──▶ agent (lane router + tools)
│
browser speaker ◀──ws── Fish Audio S2.1 TTS ◀──────┘
│ failure
▼
Fish S2.1 free via OpenRouter (same voice)
│ failure
▼
Gemini TTS (free tier)
- Gemini's server-side VAD detects the end of each utterance (live partial transcripts are shown while you talk), then the agent responds and the TTS chain speaks the reply.
- Streaming + low latency: LLM text streams in and every completed sentence
is sent to TTS while the rest is still generating. Fish runs with
FISH_LATENCY=low(~0.4 s to first synthesized audio vs ~3–5 s fornormal), so a typed message typically starts speaking in ~1.4 s end to end. - Speech formatting: every TTS engine receives sanitized text (markdown
glyphs, asterisks, and em dashes removed,
~spoken as "around"), and the system prompts instruct the models to avoid those characters in the first place. The chat bubble still shows the original text. - Fish Audio free tier: the free developer tier (
s2.1-pro-free) works with no API credit, but the model value must be sent as an HTTP header on the TTS request/WebSocket upgrade — anywhere else and Fish silently falls back to the paids2.1-pro, which returns 402 when the credit pool is empty. Vera sends the header correctly by default. - Voice overrides:
FISH_TTS_MODEL,FISH_LATENCY,OPENROUTER_TTS_MODEL,GEMINI_TTS_MODEL,GEMINI_TTS_VOICE(plus the Icarus variants) in.env. - Per-engine cooldowns: three consecutive failures sideline an engine for 60 s, and Fish direct is re-tried first every turn — topping up Fish credit instantly restores the primary voice with zero config changes.
| Engine | Limit | Reset |
|---|---|---|
Fish direct (s2.1-pro-free) |
free developer tier, fair-use (no hard cap) | — |
| Fish free via OpenRouter | 50 free-model requests/day (1000/day with $10 credit) | daily |
| Gemini TTS | ~10 requests/day | daily |
If all three are exhausted, Vera stays text-only until one resets.
Vera listens for her wake word system-level — no browser tab or mic permission needed. A small Python worker owns the microphone via ALSA and wakes the UI by broadcasting over WebSocket; saying the phrase flashes the orb and opens the mic. While the browser mic is active the worker stands down (never double listening), and an optional offline Piper voice speaks a "Yes?" prompt without blocking detection.
A custom-trained "Hey Vera" model ships in this repo at models/vera.tflite
(+ models/vera.json) and is picked up automatically on boot — no .env
change needed (or set WAKE_MODEL=models/vera.tflite explicitly). The
probability cutoff (0.96) is read from the JSON manifest, ESPHome-style, so
tuning lives with the model. It is a streaming mixednet, quantized int8 TFLite
(62 KB), trained with the microWakeWord framework from synthetic multi-speaker
samples ("Hey Vera", "Vera", command continuations) plus hard negatives of
near-homophones ("very", "verify", "every", "berry"), room reverb, and household
noise. No pretrained "Vera" model exists upstream — the stock microWakeWord
models score "Hey Vera" at ~0.02 — which is why it was trained.
The pipeline is fully scripted and regenerable end to end in ww-training/:
# full pipeline: synthetic samples -> augmentation -> features -> train
bash ww-training/pipeline.sh
# when training finishes: evaluate + install (threshold auto-sweep)
bash ww-training/finalize.sh
# probe per-clip probabilities with the installed model
python/.venv/bin/python ww-training/probe_clips.pyTo run the worker itself you need a small venv (the server spawns the worker
automatically; WAKE_ENABLED=off disables it):
python3 -m venv python/.venv
python/.venv/bin/pip install pymicro-wakeword piper-tts soundfile numpy requests
# optional, for the offline spoken "Yes?" prompt:
python/.venv/bin/piper.download_voices en_US-lessac-medium # via python -m piper.download_voices- wake (default): local streaming TFLite model — zero cloud calls, no STT
quota burn, ~100 ms reaction, <1% of a CPU core. Built-in phrases when no
custom model exists:
hey_jarvis(default),hey_mycroft,alexa,okay_nabu(WAKE_WORD). - transcript (opt-in fallback): an energy VAD waits for a completed utterance, then sends one recording to Gemini STT and wakes on a whole-word "vera" match (near-homophones like "Bira"/"Pera" count; "very"/ "verify" never match). Costs a handful of free-tier requests per hour of actual talking — never a probe every few seconds.
Tuning: raise WAKE_THRESHOLD if she false-wakes, lower it if she misses quiet
wake words; WAKE_DEVICE selects the ALSA capture device; WAKE_REFRACT sets
the seconds between wakes.
Fully offline — tests never touch the STT quota:
python/.venv/bin/python python/gen_ww_clips.py # clips → /tmp/ww-clips
python/.venv/bin/python python/eval_ww.py hey_jarvis # probability table + safe threshold window
bash scripts/run-ww-matrix.sh # whole-suite wake/no-wake matrix
python/.venv/bin/python python/wakeword_worker.py \
--mode wake --wake-word hey_jarvis --test /tmp/ww-clips/pos_hey_vera.wav # exit 0 = woke
python/.venv/bin/python scripts/test-ww-vad-offline.py # transcript-mode VAD: one probe per utteranceVera's wake engine is microWakeWord — the same framework, models, and feature pipeline ESPHome runs on ESP32-S3 boards — so the same wake phrase works on a dedicated voice satellite with identical on-device latency.
-
Hardware: any ESP32-S3 with a mic (e.g. ESP32-S3-BOX-3 or Home Assistant Voice Preview Edition).
-
Firmware: use the official ESPHome voice assistant configuration with
micro_wake_word:and the same built-in model Vera uses:micro_wake_word: models: - model: hey_jarvis # same model Vera listens for locally on_wake_word_detected: - homeassistant.service: service: rest_command.vera_wake
A custom-trained model works too — point
model:at its JSON manifest. Or skip the wake word on-device and POST to Vera from any automation:# Home Assistant configuration.yaml rest_command: vera_wake: url: http://<vera-host>:3000/api/wake method: POST
-
Vera side: nothing to configure —
POST /api/wakenotifies every open UI exactly like the spoken wake word.
MCP tools are what let Vera act: filesystem, home devices, anything you add.
The registry lives in mcp.servers.json; manage servers via the UI's MCP
overlay or REST (see docs/ADDING-MCP-SERVERS.md):
# add a stdio server at runtime
curl -X POST localhost:3000/api/mcp/servers -H 'Content-Type: application/json' \
-d '{"name":"fs","transport":"stdio","command":"npx","args":["-y","@modelcontextprotocol/server-filesystem","/home/evan"],"enabled":true}'
# list servers + discovered tools
curl localhost:3000/api/mcp/servers
# remove
curl -X DELETE localhost:3000/api/mcp/servers/fsTools are exposed to the model namespaced as mcp__<server>__<tool>.
- On boot Vera spawns
openclaw gateway(logs stream into Vera's terminal panel). Run the gateway yourself and setOPENCLAW_MODE=externalto attach instead;OPENCLAW_GATEWAY_TOKENsupports token-protected gateways. - The LLM sees a
vera_openclaw_delegatetool and can hand off multi-step jobs (research, file ops, scheduled work) autonomously; delegation works even without a live gateway via headlessopenclaw agent exec.
REST (default port 3000):
| Endpoint | Purpose |
|---|---|
POST /api/wake |
Wake all open UIs (ESP32 boards, automations) |
GET /api/lanes |
Task lanes + persona/router prompts |
GET /api/system |
System monitor snapshot + history |
GET /api/mcp/servers |
Registered servers + discovered tools |
POST /api/mcp/servers |
Add an MCP server |
DELETE /api/mcp/servers/:name |
Remove an MCP server |
WebSocket (/ws) — client → server:
| Message | Purpose |
|---|---|
mic {data} |
Base64 PCM16 16 kHz mic audio |
mic_flush |
Flush the current STT utterance |
voice_mode {on} |
Whether text-chat replies should be spoken |
mic_state {on} |
Browser mic active (pauses the wake worker) |
chat {text} |
Send a text message |
persona_toggle |
Switch Vera ⇄ Icarus |
beeps |
Debug: play three tones through the audio path |
ping |
Liveness (pong reply) |
Server → client events: transcript / transcript_partial, thinking,
lane, delta (streamed reply text, tagged with a turn id so each reply
renders as its own bubble), tool_call, openclaw_result, audio
(base64, mime-tagged), persona, wake_word, log / log_history, system.
One-off scripts in scripts/ (print results, never secrets). Provider tests
run standalone; e2e tests need a running Vera (PORT env, default 9876/9877):
node scripts/test-providers.mjs # provider connectivity (OrcaRouter + Fish WS)
node scripts/test-fish.mjs # Fish WS vs HTTP, ref/no-ref variants
node scripts/test-fish-free.mjs # proves s2.1-pro-free works: header vs start-frame
node scripts/test-fish-latency.mjs # Fish latency modes: normal vs balanced vs low
node scripts/test-gemini-node.mjs # Gemini TTS direct (3 attempts)
node scripts/test-openrouter-tts.mjs # OpenRouter TTS probes (Fish free, Deepgram free)
node scripts/test-stt-speech.mjs # full STT pipeline with synthetic speech
node scripts/test-stt-probe.mjs # force one Gemini Live STT session (2s silence)
node scripts/test-stt-variants.mjs # raw Gemini Live protocol config variants
node scripts/test-latency.mjs # end-to-end first-delta / first-audio timings
node scripts/test-e2e-voice.mjs # chat_speak → LLM reply + audio (real TTS chain)
node scripts/test-icarus-e2e.mjs # persona switch round-trip over WS (real LLM)
node scripts/test-fixes.mjs # response dedupe + chat TTS + STT regression
node scripts/test-chat-streaming.mjs # per-reply bubbles + chat layout (headless Chrome)
node scripts/test-ui.mjs # glass UI layout checks (dock, no top bar)
node scripts/test-browser-playback.mjs # real-Chrome playback (greeting + chat TTS)
node scripts/test-ww-gen.mjs # generate wake-word test clips via TTS
bash scripts/run-ww-matrix.sh # offline wake/no-wake matrix over /tmp/ww-clips
python/.venv/bin/python scripts/test-ww-vad-offline.py # transcript-mode VAD, no network
bash scripts/kiosk.sh # fullscreen kiosk launcher (finds Chromium)- No LLM replies — check the terminal overlay for
all N failover steps failed; the last provider error is shown in chat. Missing keys produce a clear warning instead. - TTS silent — see the quota table above; the
voice:counter shows what the browser received and played. - Port already in use — Vera logs the exact hint:
pkill -f "node dist/index.js". - Wake word: false wakes — raise
WAKE_THRESHOLD(e.g.0.98); missing quiet wake words — lower it a little. - Wake word: no
[wake]lines at all — the worker isn't running: check the Python venv exists andWAKE_ENABLED=on.arecord not foundmeans installalsa-utils(the worker fails fast with a clear error). - Transcript mode only:
transcript probe failed: ... Read timed outor503— Google's free STT tier is congested; probes fire once per completed utterance, so wait a moment or switch back toWAKE_MODE=wake(fully local).heard: "<noise>"means speech didn't transcribe to the wake word — adjustWAKE_ENERGY_GATE(raise in noisy rooms, lower to catch quiet speech).
- Raspberry Pi / display — see docs/DEPLOY-RPI.md for a systemd service, kiosk autostart, and performance tuning.
- Kiosk mode —
bash scripts/kiosk.shopens the UI fullscreen in Chromium with mic permissions pre-granted. - Desktop app (optional) —
npm install --save-dev electronthennpm run desktop(VERA_KIOSK=1for fullscreen,VERA_URLto point at a remote Vera).
- Personality — edit
src/prompts.ts: thePERSONASregistry (Vera and Icarus ship by default; a new assistant is a ~10-line addition) and theLANESlist (new lanes appear in the UI automatically). Rebuild withnpm run build. - Frontend — plain HTML/CSS/JS in
ui/, no build step. - More docs — docs/ARCHITECTURE.md has the module map, boot sequence, and request flow.
- Everything runs locally except the cloud AI calls you configure.
.envis gitignored — keep API keys out of git.- MCP servers run with your user's permissions; only add ones you trust.