Release status: 0.1.0-alpha. Local mode is under active stabilization. Hosted mode is development-only until authentication, isolation, quotas, and restrictive permission policy are implemented.
polycode aims to be a Claude-Code-class terminal coding agent that is provider-agnostic across OpenAI, Google Gemini, and xAI (Grok). The engine was structurally sound early on (provider-blind loop, clean sandbox/permission seams) but thin — this roadmap tracks the work to close the gap to Claude Code, phase by phase. Every item is provider-agnostic by construction: nothing in the loop, tools, or context layer is vendor-specific.
The agent was starting blind and looping without guardrails. Phase 1 gave it situational awareness and a resilient loop.
- Project-context injection — new
gatherContext()inpackages/core/src/context.tsbuilds, every session, an<environment>block (cwd, platform, date, git branch), a short<git_status>, a<directory>snapshot, and inlines project conventions when present (CLAUDE.md/AGENTS.md/POLYCODE.md/.cursorrules/.github/copilot-instructions.md, first match wins). It runs entirely through theSandboxcontract, so it behaves identically on the local and docker backends. The CLI appends it to the system prompt before launching the TUI. - Loop hardening (
packages/core/src/agent.ts) — amaxStepscap (default 50) prevents runaway tool loops; transient provider errors (rate limits, dropped connections, 5xx) retry with exponential backoff + jitter, but only before any output has streamed this turn, so text and tool calls are never duplicated. User interrupts (esc) still propagate immediately. readupgrade (packages/tools/src/index.ts) —cat -nnumbered output withoffset/limitand a continuation hint, so the model can reference and edit by line.
Verification: pnpm typecheck clean; live pnpm smoke xai:grok-4.3 round-trip green
(tool call → numbered result → text).
The tools were capable-but-thin; Phase 2 brought them to Claude-Code standard.
grep— ripgrep fast-path invoked through a no-shellexecFile(argv passed straight to the binary, so a model-supplied pattern/glob/path can't inject a command), with a much stronger JS-walk fallback supportingignore_case, aglobfilter,contextlines, andmax_results. Whenrgis not on PATH the JS path runs transparently.edituniqueness guard — errors with a match count whenold_stringis not unique (unlessreplace_all), so an ambiguous edit can't silently hit the wrong spot. A latent$-substitution bug in single-replace was fixed by switching to a literal replace.multi_edit(new) — applies a sequence of edits to one file atomically; if any edit fails, nothing is written.ls(new) — lists files and directories directly under a path.- Glob fix —
**/*.tspreviously required a path separator, so top-level files never matched; the matcher was rewritten with a proper**/globstar and?support (benefits bothgrep --globand theglobtool). - Parallel execution (
packages/core/src/agent.ts) — consecutive read-only (safe) tool calls now run concurrently viaPromise.all, with results kept in call order; mutating and dangerous tools still run sequentially so permission prompts never overlap.
Verification: pnpm typecheck clean; behavior covered by a throwaway harness (glob /
ignore-case / context grep, the uniqueness guard, atomic multi_edit, ls, parallel timing)
plus a live pnpm smoke xai:grok-4.3 round-trip. Note: ripgrep is not installed on the dev box,
so the JS grep path is the live-tested one; the rg fast-path is typecheck-verified and engages
automatically when rg is present.
Made the terminal UI feel like Claude Code rather than a raw event log.
- Edit/write diffs —
edit,multi_edit, andwriteemit a colored unified diff (LCS with collapsed context) through a UI-onlydisplaychannel onToolRunResult. The diff is shown under⎿(green+/ red-/ dim context); the model-facing tool output stays terse, so diffs never bloat the context window. - Multi-line tool output — results render several indented lines with a
… +N lineshint instead of a single truncated line. - Sticky permissions — the prompt is three-way: allow once (
y), allow for the session (a), or deny (n). The engine remembers session grants and stops re-prompting for that tool. - Token / context meter — the status bar shows
ctx <live>/<window>and a session token total, accumulated from each turn's usage and updated on model switch / route. - No-flicker rendering — finished turns are committed to an Ink
<Static>region so they print once and never repaint; only the in-progress turn and the composer redraw during streaming. Tool-call headers readRead(path)instead of raw JSON.
polycode's first on-disk surface — the conversation is now saved and resumable.
- Session transcripts — after every turn the conversation (canonical messages) is written to
<cwd>/.polycode/sessions/<id>.json(id is timestamped + random, so files sort chronologically)..polycode/is gitignored and ignored by the sandbox walk, so transcripts never pollute search/context. This is also the log/observability surface that previously didn't exist. - Resume —
--continuereopens the most recent session in the project;--resume <id>reopens a specific one. The saved messages seed the agent (AgentOptions.initialMessages) and are rebuilt into the on-screen history, so a resumed session both remembers and shows the prior conversation. A fresh system prompt (current project context) is regenerated on resume. --sessions— lists saved transcripts (id · model · title) and exits.- Layering — the
SessionStore(node:fs) lives in the CLI; the TUI stays I/O-free and just receivesinitialMessages+ anonPersistcallback. Persistence is best-effort — a write error never crashes a turn.
Verification: pnpm typecheck clean; harness round-trips a real agent conversation (save → load →
seed a fresh agent → continue), and --sessions + live smoke pass.
Phases 1–4 were "basic parity." The forward plan to Claude-Code-class — MCP, context compaction, subagents, skills/commands, web tools, hooks, checkpoints, server hardening — is laid out, sequenced, and grounded in a code audit + feature diff in the Build-out Plan. Start there.
- Phase 5.0 — pre-flight ground truth (verify model IDs, pin AI SDK, fill real context windows).
- Phase 5.1 — MCP client (force multiplier; cheap via the AI SDK's
experimental_createMCPClient). - Phase 5.2 — context compaction (correctness: history is currently unbounded).
- Phase 5.3 / 5.4 / 5.5 — subagents, skills/commands, web tools + todo.
- Agentic key provisioning — the Settings
ahook is still a stub (MCP/tool flow TODO).
| Phase | Scope | State |
|---|---|---|
| 1 | Project context · loop hardening · numbered read |
✅ done (2026-06-23) |
| 2 | Tools parity (grep/edit/ls/parallel) | ✅ done (2026-06-24) |
| 3 | TUI parity (diffs, sticky perms, token meter, no-flicker) | ✅ done (2026-06-24) |
| 4 | Persistence & logs (transcripts, --continue/--resume/--sessions) |
✅ done (2026-06-24) |