Optional web panel served by BeaconMCP on the same origin as the MCP endpoint (https://<your-host>/app/...). Three pages:
/app/login— exchanges a client id + client secret + a second factor (TOTP code or passkey) for an MCP bearer, and stores it in a 90-day HttpOnly session cookie. Removes the need to issuecurlrequests from a phone./app/chat— multi-conversation chat with Gemini 3.6 Flash / 3.5 Flash-Lite (GA) or Gemini 3.1 Pro (preview, Google allowlist required). RequiresGEMINI_API_KEY./app/tokens— generates named bearers so external MCP clients (Gemini web, ChatGPT, Assistant Desktop) can be wired up without the OAuth dance. Works withoutGEMINI_API_KEY.
The dashboard is on by default as long as a BEACONMCP_SESSION_KEY is set. Two modes:
| Mode | Condition | Active pages |
|---|---|---|
| Full | GEMINI_API_KEY set |
/app/login, /app/chat, /app/tokens |
| Tokens only | GEMINI_API_KEY absent |
/app/login, /app/tokens (chat redirects to tokens) |
- (optional) Add a Gemini API key to enable the integrated chat:
GEMINI_API_KEY=...
- The session encryption key (
BEACONMCP_SESSION_KEY) is generated byinstall.shon first run. For manual deployments:echo "BEACONMCP_SESSION_KEY=$(openssl rand -base64 32)" >> /opt/beaconmcp/.env
- Restart:
systemctl restart beaconmcp.
At boot, the server logs the active dashboard URL and whether the chat is enabled:
Dashboard: http://0.0.0.0:8420/app/login (chat: enabled)
or (chat: disabled, tokens only) when the Gemini key is missing, or disabled when BEACONMCP_DASHBOARD_ENABLED=false.
- Navigate to
https://<your-host>/on any device — the root path redirects to/app/login. - Enter client id + client secret, then the current TOTP code — or use a passkey instead, from the link under the code boxes (see Passkeys). Enter validates as soon as the sixth digit lands.
- A confirmation screen shows when the MCP access and the signed-in session expire, offers to register a passkey on this device, and waits for Finish signing in.
- Land on
/app/chat(or/app/tokensin tokens-only mode) with the conversation history restored. - Every 24 hours the underlying MCP bearer expires. The dashboard prompts only for a fresh TOTP code on
/app/refresh; the client id and secret remain encrypted server-side.
With JavaScript disabled the page falls back to the previous behaviour: a plain form POST that redirects straight to the landing page, no confirmation screen and no passkeys.
A passkey replaces the 6-digit code, not the client secret — signing in still needs both factors. Enrol one from the confirmation screen right after a normal sign-in, then on the next visit use Use a passkey instead under the code boxes.
The same option is on the OAuth consent page (/oauth/authorize), which now also confirms before
redirecting: it shows when the granted access expires, offers passkey enrolment, and only mints the
authorization code when you press Finish signing in.
Registered passkeys are listed at the bottom of /app/tokens with their label, when they were added
and when they were last used, each with a Remove button.
Caveats — hostname binding, the HTTPS/loopback requirement, and why you should keep your TOTP seed — are covered in security.md.
Reached from the sidebar or directly at /app/tokens. Intended for users wiring BeaconMCP into an external MCP client rather than using the integrated chat.
It exposes:
- The MCP URL to paste into the external client (with a copy button).
- A creation form that requires a name (60 characters max, e.g. "Gemini Web", "ChatGPT macOS") plus the current TOTP code.
- The newly generated token, shown once in an orange card with a copy button. Reloading the page removes it from view.
- The list of active tokens: name, 12-character prefix, hours until expiry, revoke button.
Constraints:
| Setting | Value |
|---|---|
| Expiry | 24 h (inherits TokenStore.TOKEN_TTL) |
| Per-client cap | 3 active tokens |
| TOTP | Re-verified on every creation |
| Revocation | By prefix (6 chars minimum), scoped to the owning client |
| Storage | In-memory — systemctl restart invalidates every token |
- Gemini 3.6 Flash — GA, available on every AI Studio key. Used by default (
gemini-3.6-flash). - Gemini 3.5 Flash-Lite — GA, the cheap high-throughput option (
gemini-3.5-flash-lite), 5× cheaper in and out than 3.6 Flash. - Gemini 3.1 Pro (preview) — gated by a Google allowlist. Without allowlist access, the dashboard surfaces a clear message pointing back to the two GA models.
- Thinking effort — dropdown with
minimal/low/medium/high. The Gemini 3 family takes athinking_levelenum; the token-budget mapping in_BUDGET_BY_EFFORT(and the 128-token floorgemini-2.5-proneeded) only applies to models outside it, and is kept for conversations that predate the switch.
Gemini 2.5 Flash / Pro and gemini-3-flash-preview were retired from the picker when 3.6 Flash and 3.5 Flash-Lite went GA (2026-07-21). A conversation sitting on one of them is moved to the closest current model by the schema-6 migration; messages.model is left alone, so the transcript keeps naming whichever model actually wrote each reply, and usage.py keeps their rates so old turns are never re-priced.
- Markdown rendering — the client parses headings (
#–######), ordered and unordered lists, blockquotes, horizontal rules, code fences (withlang-*class), inline code, bold/italic/strikethrough, and HTTP(S) links.
The model's input is untrusted: a log line, a config file, or a web-search result can carry an injected instruction. So two classes of tool never run without manual approval from the chat UI.
Arbitrary code execution — ssh_run, proxmox_run, and the primitives that become code execution in one hop (proxmox_write_file, proxmox_upload_file, proxmox_download_file, proxmox_delete_transfer). Writing ~/.ssh/authorized_keys or a file under /etc/cron.d is exactly as good as a shell, which is why the file tools sit alongside the exec tools.
Destructive or irreversible — vm_bulk_action, proxmox_vm_stop, proxmox_vm_restart, proxmox_vm_migrate, proxmox_snapshot_rollback, proxmox_snapshot_delete, proxmox_backup_restore, bmc_power_off, bmc_power_reset, and beaconmcp_self_update with confirm=True — which pulls new code, reinstalls dependencies and restarts the service, i.e. replaces the very process enforcing this gate.
Three call shapes are let through without a modal, because they don't change anything:
ssh_run/proxmox_runcarrying onlyexec_id=— that's read-only polling of an already-approved session.proxmox_snapshot_create/_rollback/_deletecalled withdry_run=True— they only report what they would do. The exemption is limited to those three by name, never inferred from the argument: an undeclareddry_runis silently dropped during argument validation, so trusting it would letssh_run(command=..., dry_run=True)past the modal and then run for real.proxmox_vm_configwithoutupdates, andbeaconmcp_self_updatewithoutconfirm=True— the read shape of a read-or-write tool. Reading the argument is sound for these two because both parameters are declared by the tool, so what the gate reads is what the tool acts on; that is precisely what makes it unsound for an undeclareddry_run.
When Gemini fires a gated call:
- The tool card switches to an "approval required" state (orange badge, auto-expanded so arguments are visible).
- Two buttons: Approve / Reject.
- The Gemini turn blocks server-side until the decision is made (5-minute timeout).
- On rejection, Gemini receives a
FunctionResponse {"error": "user_rejected"}and can revise its reply.
The allow-list is hard-coded in src/beaconmcp/dashboard/chat.py (_NEEDS_CONFIRMATION, _CONFIRM_WHEN_ARG_PRESENT, _tool_call_requires_confirmation). Only the integrated chat enforces this gate; external MCP clients (Assistant Desktop, Gemini CLI, ChatGPT MCP) must enable their own per-call approval mode (see the Security section of the root README).
Tools that carry _meta.ui.resourceUri — proxmox_vm_panel, proxmox_logs_panel, cluster_overview_interactive — render as a live interface in the chat instead of a block of JSON. The dashboard implements both halves of the MCP Apps extension: it announces io.modelcontextprotocol/ui when it opens its MCP session, and it plays host to the ui:// document over postMessage.
The extension is declared through ClientCapabilities.extensions, a field mcp only types in 2.0. The <2 pin is not in the way: the model accepts extra fields, so the capability serialises under the name the spec gives it and the server reads the same JSON either way (dashboard/mcp_bridge.py).
The document is served by /app/api/mcp/panel and framed with sandbox="allow-scripts" and no allow-same-origin. That puts it on an opaque origin: it cannot read the session cookie, cannot read the CSRF token, cannot reach into the parent page. Its response also carries its own Content-Security-Policy — default-src 'none', connect-src 'none', frame-ancestors 'self' — so it cannot open a socket of its own either.
What is left is postMessage to the parent. Every tool call a panel makes therefore goes through /app/api/mcp/call, which is session-authenticated and CSRF-protected, and where the policy below is applied.
A panel button is a labelled control a human clicked, so the approval modal — which exists because the model's input is untrusted — would restate the click rather than check it. Panel calls are therefore not gated. What is not granted is a blanket exemption for anything running in a frame: a ui:// document is HTML the server wrote, and this dashboard is a general MCP host, so a blanket rule would hand every connected server a way around the gate it is documented to be subject to.
The exemption is a closed list, enforced server-side in panel_call_allowed():
proxmox_vm_start/_stop/_restart— one guest per call, visible in the panel, reversible from it.proxmox_vm_config, but only when every key inupdatesis sizing (cores,sockets,memory,balloon,cpulimit,cpuunits). Exempting the tool itself would exempthookscript, raw QEMUargsand device passthrough along with it.- Everything that was never gated in the first place — the read-only tools, including the three panel tools themselves.
Anything else is refused with 403 confirmation_required, and the panel shows the reason. It is refused rather than prompted because there is no turn in flight to hang a modal on — and because the panel already has a way through: ui/message hands the request to the model, which puts it back under the modal where it belongs.
A panel that acts on the cluster pushes the fresh state back with ui/update-model-context. The page holds the latest update per panel and sends it with the next message, labelled as coming from the panel rather than from the operator. Without it, stopping a VM from the panel would leave the next turn believing it still runs — the button's result goes to the iframe, not into the conversation.
Only the ui:// URI is stored with the tool call, never the snapshot behind it. A panel in an older conversation renders as an Open panel button; clicking it mounts the frame and refetches. Live figures in a panel that has been sitting in the transcript for a week would be worse than a short spinner.
SQLite at /opt/beaconmcp/dashboard.db (WAL mode). Five tables:
sessions— cookie → client id + AES-GCM-encrypted client secret + current bearer.conversations— title, model, effort, owning client.messages— user/assistant, content, tool calls JSON, thinking summary.usage_events— per-turn ledger: client id, tokens (prompt/cached/output), USD cost, timestamp.usage_5h_sessions— one row per client holding the current 5-hour session (materialized to avoid aGROUP BYon every pre-check).
Everything is scoped by client_id; a single client can hold multiple active sessions (phone + laptop).
Each chat turn computes its USD cost from the usage_metadata returned by Gemini. Input tokens are billed at the cached-discount rate whenever cachedContentTokenCount is non-zero — Gemini 2.5+ applies implicit caching automatically once the prompt crosses 1024 tokens for Flash or 4096 tokens for Pro, with no client-side work.
Two windows are enforced per OAuth client:
| Window | Semantics | Env variable | Default |
|---|---|---|---|
| 5h | Anthropic-style session: opens on the first message after ≥5 h of idle time, lasts exactly 5 h, then closes. | BEACONMCP_DASHBOARD_LIMIT_5H_USD |
2.0 |
| Week | Rolling sum over the last 7 days. | BEACONMCP_DASHBOARD_LIMIT_WEEK_USD |
10.0 |
Setting a variable to 0 disables that window. When a cap is exceeded, the next request is rejected before being sent to Gemini, with a message stating when the 5 h window resets. (When beaconmcp.yaml is used, the same caps are configured under features.dashboard.limits.)
The chat footer shows a compact 5H XX% · 7D XX% line updated after every turn via an SSE usage_update event. Clicking the bar opens a modal with progress bars, the 5 h reset time, a "rolling 7-day window" label, and a refresh button.
Rates used (USD per 1 M tokens, aligned with the public Google AI Studio pricing on 2026-07-29). The retired rows are kept because the ledger re-prices stored turns, and dropping a rate would silently re-bill that history at the fallback model's price:
| Model | Input | Cached | Output |
|---|---|---|---|
gemini-3.6-flash |
$1.50 | $0.15 | $7.50 |
gemini-3.5-flash-lite |
$0.30 | $0.03 | $2.50 |
gemini-3.1-pro-preview (≤200k / >200k) |
$2.00 / $4.00 | $0.20 / $0.40 | $12.00 / $18.00 |
gemini-2.5-flash (retired) |
$0.30 | $0.03 | $2.50 |
gemini-2.5-pro (retired, ≤200k / >200k) |
$1.25 / $2.50 | $0.125 / $0.25 | $10.00 / $15.00 |
gemini-3-flash-preview (retired) |
$0.50 | $0.05 | $3.00 |
Constants live in src/beaconmcp/dashboard/usage.py — update them when Google adjusts its prices.
The dashboard keeps its own MCP session (streamablehttp_client + ClientSession) pointed at the local /mcp endpoint (http://127.0.0.1:8420/mcp). Tools are converted manually into FunctionDeclaration objects and the function_call / function_response loop is orchestrated server-side (AFC SDK disabled via AutomaticFunctionCallingConfig(disable=True)). This works around known google-genai bugs with Gemini 2.5 Pro + MCP + streaming + thinking.
A remote mode (server-driven
McpServer) exists but is disabled: requests through Cloudflare Tunnel lose theAuthorizationheader, producing systematic 500 INTERNAL responses. SettingBEACONMCP_DASHBOARD_MCP_MODE=remotelogs a startup warning and returns an actionable error on every chat turn.
TokenStore lives in memory. After systemctl restart beaconmcp, bearers are invalidated while dashboard sessions (SQLite) persist. The dashboard detects this by calling TokenStore.validate() on every sensitive route; when a bearer is gone but the session timestamp is still valid, the user is routed to /app/refresh to enter a fresh TOTP code and mint a new bearer.
Consequence for externally-issued tokens (/app/tokens): a service restart forces every Gemini-web / ChatGPT / Assistant-Desktop integration to regenerate its token. If this is operationally annoying, move TokenStore to SQLite (not done today).
BEACONMCP_DASHBOARD_ENABLED=falseTo keep the tokens page but drop the chat, leave GEMINI_API_KEY unset.