Skip to content

Repository files navigation

Vera — Local AI Home Assistant

Vera is a local-first AI home assistant with voice, wake word detection, MCP tools, OpenClaw integration, and a polished browser UI. It is designed to run on a Raspberry Pi or Mini-PC and speaks with streaming, low-latency TTS.

What's inside

Piece Tech
Core server Node.js + TypeScript (ESM), Express, ws
LLM Gemini gemini-3.5-flash-lite (primary) → OrcaRouter → OpenRouter failover chain
Voice in Gemini Live streaming transcription (gemini-3.5-transcribe-live)
Voice out Fish Audio S2.1 → Fish S2.1 free via OpenRouter → Gemini TTS (auto failover)
Tools MCP servers (stdio + HTTP) via @modelcontextprotocol/sdk
Automation OpenClaw gateway (spawned by Vera) + headless task delegation
Wake word Headless Python worker, local microWakeWord TFLite model (trained "Hey Vera" model included)
UI Vanilla JS/HTML/CSS — chat, terminal, avatar, system monitor, MCP manager (no build step)
Desktop Optional Electron wrapper (npm run desktop) or Chromium kiosk (scripts/kiosk.sh)

Quick start

# 1. Install (Node >= 22.22.3)
npm install

# 2. Configure
cp .env.example .env
# edit .env — see "API keys" below for what is actually required

# 3. Run
npm start               # build + serve on http://localhost:3000
npm run dev             # or: auto-reloading dev server (tsx watch)
npm run typecheck       # typecheck without emitting

On boot Vera starts the OpenClaw gateway, the MCP registry servers, the system monitor, and (unless disabled) the wake word worker. Open the UI at http://localhost:3000 — the greeting is spoken once the browser sends its first message.

API keys

Everything runs locally except the cloud AI calls you configure. Keys live in .env (never commit it):

Key Needed for
ORCAROUTER_API_KEY or OPENROUTER_API_KEY Text chat (at least one — these enable the LLM)
GEMINI_API_KEY Primary chat model, all speech recognition (STT), and fallback TTS
FISH_API_KEY Primary TTS voice (Fish Audio S2.1); free tier works

Text chat works with just an OrcaRouter/OpenRouter key. Voice needs Gemini for speech-to-text; the speaking voice then fails over Fish → OpenRouter → Gemini automatically, so a single GEMINI_API_KEY plus one LLM key is a working setup.

Other options

.env.example documents every variable. The ones you'll most likely touch:

  • PORT — UI/server port (default 3000).
  • FISH_REFERENCE_ID / FISH_ICARUS_REFERENCE_ID — custom cloned voices for Vera / Icarus.
  • WAKE_ENABLED, WAKE_MODE, WAKE_WORD, WAKE_MODEL, WAKE_THRESHOLD — wake word (below).
  • OPENCLAW_MODE — spawn (Vera launches the gateway, default) or external (attach to a gateway you started yourself with openclaw gateway).

Using Vera

The UI

Glass theme: pure black background, translucent panels with glowing white outlines, gold accents (user messages, speaking orb, live mic).

  • Main screen — the talking orb, chat history, system monitor strip, and a floating translucent dock at the bottom. No top bar, no tabs.
  • Dock — persona toggle (Vera ⇄ Icarus), voice output toggle, mic button, chat view, terminal overlay, system monitor toggle, MCP servers overlay, send.
  • Terminal — mirrors every server log line live (recent history is replayed to new clients).
  • MCP — list / add / remove MCP servers at runtime.

Text chat works with or without voice. The mic button opens the browser mic; Vera speaks replies through a streaming TTS chain, and every spoken frame is counted by the voice: counter in the bottom-right of the UI.

Personas

Two assistants ship by default, each with her/his own prompt, voice, theme, and conversation context: Vera (warm, witty) and Icarus (dry, precise). Switch by saying "Icarus mode" / "Vera mode" / "Icarus off", or with the dock toggle. A switch resets context, relabels the chat, and is announced with a spoken greeting. Give Icarus his own voice with FISH_ICARUS_REFERENCE_ID and GEMINI_ICARUS_TTS_VOICE.

Task lanes

Every message is routed into a lane that shapes the reply: Quick, Smart Home, Coding, Business, Research. A regex heuristic routes most messages instantly; ambiguous ones go through an LLM router prompt. The Quick lane runs with no tools attached (tool definitions measurably slow fast models' first token from ~0.6s to ~13s); other lanes can call MCP tools and OpenClaw delegation, for up to 3 tool rounds per turn. When the model goes straight to work, you get a spoken "on it" acknowledgement first, then the answer. Conversation history is bounded to the last 40 turns per connection.

Model failover chain

Requests walk this chain in order; the first provider that streams wins:

  1. gemini/gemini-3.5-flash-lite — primary (needs GEMINI_API_KEY, fast, free tier)
  2. orcarouter/deepseek/deepseek-v4-flash-free
  3. orcarouter/free
  4. Steps 2–3 again after a 4 s delay (usage-limit backoff)
  5. openrouter/minimax/minimax-m2.7 (cheap paid — the old :free slug was removed upstream)
  6. openrouter/free (OpenRouter's auto free router)

If all steps fail, the error is logged and shown in chat.

Failover is aggressive and visible in the terminal panel:

  • Rate-limited providers are skipped for up to 30 minutes (honoring the provider's own retry-after hint when it sends one) instead of being retried mid-conversation.
  • Model slugs that 404 or are "unavailable for free" are skipped for an hour.
  • A provider that accepts the request but never streams is cooled down for 60 s.
  • Streams are protected by hang watchdogs: first chunk must arrive within 12 s (30 s when tools are attached), chunks must keep flowing (15 s stall limit), and the total stream is capped at 180 s.
  • The OpenAI SDK's internal retries are disabled so failover happens here, visibly, in the log.

Voice pipeline

browser mic (PCM16 16 kHz) ──ws──▶ Gemini Live STT ──▶ agent (lane router + tools)
                                                           │
        browser speaker ◀──ws── Fish Audio S2.1 TTS ◀──────┘
                                   │ failure
                                   ▼
                    Fish S2.1 free via OpenRouter (same voice)
                                   │ failure
                                   ▼
                            Gemini TTS (free tier)
  • Gemini's server-side VAD detects the end of each utterance (live partial transcripts are shown while you talk), then the agent responds and the TTS chain speaks the reply.
  • Streaming + low latency: LLM text streams in and every completed sentence is sent to TTS while the rest is still generating. Fish runs with FISH_LATENCY=low (~0.4 s to first synthesized audio vs ~3–5 s for normal), so a typed message typically starts speaking in ~1.4 s end to end.
  • Speech formatting: every TTS engine receives sanitized text (markdown glyphs, asterisks, and em dashes removed, ~ spoken as "around"), and the system prompts instruct the models to avoid those characters in the first place. The chat bubble still shows the original text.
  • Fish Audio free tier: the free developer tier (s2.1-pro-free) works with no API credit, but the model value must be sent as an HTTP header on the TTS request/WebSocket upgrade — anywhere else and Fish silently falls back to the paid s2.1-pro, which returns 402 when the credit pool is empty. Vera sends the header correctly by default.
  • Voice overrides: FISH_TTS_MODEL, FISH_LATENCY, OPENROUTER_TTS_MODEL, GEMINI_TTS_MODEL, GEMINI_TTS_VOICE (plus the Icarus variants) in .env.
  • Per-engine cooldowns: three consecutive failures sideline an engine for 60 s, and Fish direct is re-tried first every turn — topping up Fish credit instantly restores the primary voice with zero config changes.

Voice quota limits (why TTS can go silent)

Engine Limit Reset
Fish direct (s2.1-pro-free) free developer tier, fair-use (no hard cap) —
Fish free via OpenRouter 50 free-model requests/day (1000/day with $10 credit) daily
Gemini TTS ~10 requests/day daily

If all three are exhausted, Vera stays text-only until one resets.

Wake word (local, headless)

Vera listens for her wake word system-level — no browser tab or mic permission needed. A small Python worker owns the microphone via ALSA and wakes the UI by broadcasting over WebSocket; saying the phrase flashes the orb and opens the mic. While the browser mic is active the worker stands down (never double listening), and an optional offline Piper voice speaks a "Yes?" prompt without blocking detection.

The trained "Hey Vera" model

A custom-trained "Hey Vera" model ships in this repo at models/vera.tflite (+ models/vera.json) and is picked up automatically on boot — no .env change needed (or set WAKE_MODEL=models/vera.tflite explicitly). The probability cutoff (0.96) is read from the JSON manifest, ESPHome-style, so tuning lives with the model. It is a streaming mixednet, quantized int8 TFLite (62 KB), trained with the microWakeWord framework from synthetic multi-speaker samples ("Hey Vera", "Vera", command continuations) plus hard negatives of near-homophones ("very", "verify", "every", "berry"), room reverb, and household noise. No pretrained "Vera" model exists upstream — the stock microWakeWord models score "Hey Vera" at ~0.02 — which is why it was trained.

The pipeline is fully scripted and regenerable end to end in ww-training/:

# full pipeline: synthetic samples -> augmentation -> features -> train
bash ww-training/pipeline.sh
# when training finishes: evaluate + install (threshold auto-sweep)
bash ww-training/finalize.sh
# probe per-clip probabilities with the installed model
python/.venv/bin/python ww-training/probe_clips.py

To run the worker itself you need a small venv (the server spawns the worker automatically; WAKE_ENABLED=off disables it):

python3 -m venv python/.venv
python/.venv/bin/pip install pymicro-wakeword piper-tts soundfile numpy requests
# optional, for the offline spoken "Yes?" prompt:
python/.venv/bin/piper.download_voices en_US-lessac-medium   # via python -m piper.download_voices

Detection modes (WAKE_MODE)

  1. wake (default): local streaming TFLite model — zero cloud calls, no STT quota burn, ~100 ms reaction, <1% of a CPU core. Built-in phrases when no custom model exists: hey_jarvis (default), hey_mycroft, alexa, okay_nabu (WAKE_WORD).
  2. transcript (opt-in fallback): an energy VAD waits for a completed utterance, then sends one recording to Gemini STT and wakes on a whole-word "vera" match (near-homophones like "Bira"/"Pera" count; "very"/ "verify" never match). Costs a handful of free-tier requests per hour of actual talking — never a probe every few seconds.

Tuning: raise WAKE_THRESHOLD if she false-wakes, lower it if she misses quiet wake words; WAKE_DEVICE selects the ALSA capture device; WAKE_REFRACT sets the seconds between wakes.

Test the wake word without speaking

Fully offline — tests never touch the STT quota:

python/.venv/bin/python python/gen_ww_clips.py            # clips → /tmp/ww-clips
python/.venv/bin/python python/eval_ww.py hey_jarvis      # probability table + safe threshold window
bash scripts/run-ww-matrix.sh                             # whole-suite wake/no-wake matrix
python/.venv/bin/python python/wakeword_worker.py \
  --mode wake --wake-word hey_jarvis --test /tmp/ww-clips/pos_hey_vera.wav   # exit 0 = woke
python/.venv/bin/python scripts/test-ww-vad-offline.py    # transcript-mode VAD: one probe per utterance

ESP32 / ESPHome

Vera's wake engine is microWakeWord — the same framework, models, and feature pipeline ESPHome runs on ESP32-S3 boards — so the same wake phrase works on a dedicated voice satellite with identical on-device latency.

  1. Hardware: any ESP32-S3 with a mic (e.g. ESP32-S3-BOX-3 or Home Assistant Voice Preview Edition).

  2. Firmware: use the official ESPHome voice assistant configuration with micro_wake_word: and the same built-in model Vera uses:

    micro_wake_word:
      models:
        - model: hey_jarvis        # same model Vera listens for locally
      on_wake_word_detected:
        - homeassistant.service:
            service: rest_command.vera_wake

    A custom-trained model works too — point model: at its JSON manifest. Or skip the wake word on-device and POST to Vera from any automation:

    # Home Assistant configuration.yaml
    rest_command:
      vera_wake:
        url: http://<vera-host>:3000/api/wake
        method: POST
  3. Vera side: nothing to configure — POST /api/wake notifies every open UI exactly like the spoken wake word.

MCP servers

MCP tools are what let Vera act: filesystem, home devices, anything you add. The registry lives in mcp.servers.json; manage servers via the UI's MCP overlay or REST (see docs/ADDING-MCP-SERVERS.md):

# add a stdio server at runtime
curl -X POST localhost:3000/api/mcp/servers -H 'Content-Type: application/json' \
  -d '{"name":"fs","transport":"stdio","command":"npx","args":["-y","@modelcontextprotocol/server-filesystem","/home/evan"],"enabled":true}'

# list servers + discovered tools
curl localhost:3000/api/mcp/servers

# remove
curl -X DELETE localhost:3000/api/mcp/servers/fs

Tools are exposed to the model namespaced as mcp__<server>__<tool>.

OpenClaw integration

  • On boot Vera spawns openclaw gateway (logs stream into Vera's terminal panel). Run the gateway yourself and set OPENCLAW_MODE=external to attach instead; OPENCLAW_GATEWAY_TOKEN supports token-protected gateways.
  • The LLM sees a vera_openclaw_delegate tool and can hand off multi-step jobs (research, file ops, scheduled work) autonomously; delegation works even without a live gateway via headless openclaw agent exec.

HTTP + WebSocket API

REST (default port 3000):

Endpoint Purpose
POST /api/wake Wake all open UIs (ESP32 boards, automations)
GET /api/lanes Task lanes + persona/router prompts
GET /api/system System monitor snapshot + history
GET /api/mcp/servers Registered servers + discovered tools
POST /api/mcp/servers Add an MCP server
DELETE /api/mcp/servers/:name Remove an MCP server

WebSocket (/ws) — client → server:

Message Purpose
mic {data} Base64 PCM16 16 kHz mic audio
mic_flush Flush the current STT utterance
voice_mode {on} Whether text-chat replies should be spoken
mic_state {on} Browser mic active (pauses the wake worker)
chat {text} Send a text message
persona_toggle Switch Vera ⇄ Icarus
beeps Debug: play three tones through the audio path
ping Liveness (pong reply)

Server → client events: transcript / transcript_partial, thinking, lane, delta (streamed reply text, tagged with a turn id so each reply renders as its own bubble), tool_call, openclaw_result, audio (base64, mime-tagged), persona, wake_word, log / log_history, system.

Diagnostics

One-off scripts in scripts/ (print results, never secrets). Provider tests run standalone; e2e tests need a running Vera (PORT env, default 9876/9877):

node scripts/test-providers.mjs        # provider connectivity (OrcaRouter + Fish WS)
node scripts/test-fish.mjs             # Fish WS vs HTTP, ref/no-ref variants
node scripts/test-fish-free.mjs        # proves s2.1-pro-free works: header vs start-frame
node scripts/test-fish-latency.mjs     # Fish latency modes: normal vs balanced vs low
node scripts/test-gemini-node.mjs      # Gemini TTS direct (3 attempts)
node scripts/test-openrouter-tts.mjs   # OpenRouter TTS probes (Fish free, Deepgram free)
node scripts/test-stt-speech.mjs       # full STT pipeline with synthetic speech
node scripts/test-stt-probe.mjs        # force one Gemini Live STT session (2s silence)
node scripts/test-stt-variants.mjs     # raw Gemini Live protocol config variants
node scripts/test-latency.mjs          # end-to-end first-delta / first-audio timings
node scripts/test-e2e-voice.mjs        # chat_speak → LLM reply + audio (real TTS chain)
node scripts/test-icarus-e2e.mjs       # persona switch round-trip over WS (real LLM)
node scripts/test-fixes.mjs            # response dedupe + chat TTS + STT regression
node scripts/test-chat-streaming.mjs   # per-reply bubbles + chat layout (headless Chrome)
node scripts/test-ui.mjs               # glass UI layout checks (dock, no top bar)
node scripts/test-browser-playback.mjs # real-Chrome playback (greeting + chat TTS)
node scripts/test-ww-gen.mjs           # generate wake-word test clips via TTS
bash scripts/run-ww-matrix.sh          # offline wake/no-wake matrix over /tmp/ww-clips
python/.venv/bin/python scripts/test-ww-vad-offline.py  # transcript-mode VAD, no network
bash scripts/kiosk.sh                  # fullscreen kiosk launcher (finds Chromium)

Troubleshooting

  • No LLM replies — check the terminal overlay for all N failover steps failed; the last provider error is shown in chat. Missing keys produce a clear warning instead.
  • TTS silent — see the quota table above; the voice: counter shows what the browser received and played.
  • Port already in use — Vera logs the exact hint: pkill -f "node dist/index.js".
  • Wake word: false wakes — raise WAKE_THRESHOLD (e.g. 0.98); missing quiet wake words — lower it a little.
  • Wake word: no [wake] lines at all — the worker isn't running: check the Python venv exists and WAKE_ENABLED=on. arecord not found means install alsa-utils (the worker fails fast with a clear error).
  • Transcript mode only: transcript probe failed: ... Read timed out or 503 — Google's free STT tier is congested; probes fire once per completed utterance, so wait a moment or switch back to WAKE_MODE=wake (fully local). heard: "<noise>" means speech didn't transcribe to the wake word — adjust WAKE_ENERGY_GATE (raise in noisy rooms, lower to catch quiet speech).

Deploying

  • Raspberry Pi / display — see docs/DEPLOY-RPI.md for a systemd service, kiosk autostart, and performance tuning.
  • Kiosk mode — bash scripts/kiosk.sh opens the UI fullscreen in Chromium with mic permissions pre-granted.
  • Desktop app (optional) — npm install --save-dev electron then npm run desktop (VERA_KIOSK=1 for fullscreen, VERA_URL to point at a remote Vera).

Making it yours

  • Personality — edit src/prompts.ts: the PERSONAS registry (Vera and Icarus ship by default; a new assistant is a ~10-line addition) and the LANES list (new lanes appear in the UI automatically). Rebuild with npm run build.
  • Frontend — plain HTML/CSS/JS in ui/, no build step.
  • More docs — docs/ARCHITECTURE.md has the module map, boot sequence, and request flow.

Safety notes

  • Everything runs locally except the cloud AI calls you configure.
  • .env is gitignored — keep API keys out of git.
  • MCP servers run with your user's permissions; only add ones you trust.

About

A smart home/personal assistant coded in Typescript, Javascript, and Python. Uses gemini free tier, openrouter, and orcarouter. Uses Fish.Audio for audio output.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages