Skip to content

Track agent-browser + vibium as managed components — visibility, not a fix request #189

Description

@lafinak

Track agent-browser (and vibium) as managed components — visibility, not a fix request

TL;DR

ak doesn't know about the Chrome-for-Testing binaries that agent-browser and vibium
download and cache, even though it already tracks equivalent caches for Playwright and
Puppeteer. This isn't a request to fix anything — both tools are standalone (not
components of ruflo), so per your own framing below, this is squarely an "include it so
it's under management and reporting" ask. Real numbers and a real ak system --deep run
included below, not hypotheticals.

Also surfaced along the way: ruflo itself already has a design (ADR-122, Proposed
2026-05-18) to ship exactly this kind of agent-browser health-reporting via
ruflo doctor — it just hasn't landed yet, four months later. Worth knowing about before
deciding where this should live.

Your framing (quoting you back, so we're aligned)

Keep me honest why don't you log a GitHub issue in my repo... I would like a little bit
more detail like please explain what it is that you're looking for.

I don't necessarily fix problems. I report them upstream. And then when things are
working, I include them.

If this is just about including the component so it's under management and reporting on
its health, I think I can do that.

If it's a component of ruflo then we have to wait for it to have a fix. If it's
something that stands on its own, maybe there's something we can do.

Both tools involved here stand on their own — confirmed by reading each package's own
package.json:

Tool Owner Repo
agent-browser Vercel Labs github.com/vercel-labs/agent-browser
vibium VibiumDev github.com/VibiumDev/vibium — vibium.com

Neither is a ruflo component. Per your own rule, this is the "maybe there's something we
can do" case, not the "wait for ruflo to fix" case.

How this started (the actual bug, already filed elsewhere)

While using ruflo-browser's MCP tools (mcp__ruflo__browser_* — the 23-tool plugin,
more on that below) for real browser automation, agent-browser's own Chrome-for-Testing
download (~194MB) reliably failed at ~95%, 3 retries, same point every time, silently
exiting 0 despite failing. Root-caused this properly (not guessed):

  • A plain curl to the exact same URL completed cleanly in 124s at ~1.56MB/s.
  • agent-browser's own binary embeds AGENT_BROWSER_DEFAULT_TIMEOUT /
    AGENT_BROWSER_IDLE_TIMEOUT_MS env knobs (confirmed via strings on the Rust binary —
    it's reqwest + tokio::time::timeout).
  • Initially attributed the failure to a too-tight fixed timeout. Later corrected: after
    fixing an unrelated full network outage in the same sandbox (a container restart was
    required — DNS resolved fine, but every outbound TCP connection was timing out), re-ran
    the exact same agent-browser install with no override and it succeeded cleanly, 100%,
    1m42s, no retries. So the real explanation is more likely a transient network
    degradation at the time of the original failures, not a permanently-too-tight timeout.

Filed with full detail, later corrected in a follow-up comment once the network issue was
understood: ciprianmelian/ruflo-aqe-kit#21
(ciprianmelian/ruflo-aqe-kit#21). That's the right home for the
bug itself — this issue here is only about ak visibility, not about re-litigating the
bug.

The concrete gap, with real numbers from this machine

Workaround while debugging: agent-browser was pointed at an already-downloaded,
same-version Chrome binary that a completely unrelated tool (vibium, via a
qe-browser skill) had cached — ~/.cache/vibium/chrome-for-testing/152.0.7977.64/.
Once the network issue was separately fixed, agent-browser also completed its own real
download successfully — so this machine now has two full, duplicate Chrome-for-Testing
152.0.7977.64 installs
:

413M  ~/.cache/vibium/chrome-for-testing/152.0.7977.64
391M  ~/.agent-browser/browsers/chrome-152.0.7977.64

~800MB of duplicate binary. Ran ak system --deep on this exact machine to check whether
ak already sees this — full, real, unedited output:

Storage  — measured 2026-09-01 02:16
  total           934 MB · 4,028 files

  CATEGORY          SIZE     FILES
  learning-stores   886 MB   3,826
  transcripts       45.2 MB  95
  kit-caches        1.85 MB  49
  ledgers-and-logs  1.23 MB  58

  reclaimable (advisory only — ak system removes nothing)
  CANDIDATE                                            SIZE     CLEANUP
  npm content-addressable cache                        1.65 GB  npm cache clean --force
  orphaned worktree record "pre-gastown-ruflo-repair"  0 B      git worktree prune

Four storage categories, two reclaimable candidates. Zero mention of either Chrome cache,
anywhere. ak system's Install table already tracks ruflo, agentic-qe, Claude Code,
OpenCode, and RuvNet Brain (version / install method / size) — this is asking for
agent-browser and vibium to join that same table, using the same shape you already
have.

What I'm actually asking for

Two things, both purely additive to the existing ak system model:

  1. Install-table entry for agent-browser (and vibium, since it's already a
    dependency this stack pulls in via the qe-browser skill) — version, size, install
    method, same columns as the existing rows.
  2. A health signal, not just presence. agent-browser doctor already reports whether
    it has a working Chrome binary:
    Chrome
      fail  No Chrome binary found
      pass  Chrome for Testing CDN reachable (280ms, HTTP 200)
      fail  Browser launch failed: Chrome not found. Checked: ...
    
    Surfacing that signal in ak system would have caught the original problem in feat(providers): detect + configure claude/codex hosts and LLM providers #21
    immediately — instead of requiring a manual filesystem investigation across multiple
    cache directories to even discover a working binary existed.

Not asking for a fix to the download bug itself — that's Vercel Labs' code, outside your
control or ciprianmelian's, and it's already filed at the right place.

One more thing worth knowing before deciding where this lives

While digging into ruflo-browser, found ruflo/v3/docs/adr/ADR-122-browser-beyond-sota.md
(Status: Proposed, dated 2026-05-18, still Proposed as of today — ruflo doctor
still says nothing about agent-browser, confirmed by actually running it). It documents
that ruflo has two separate, drifting browser systems internally:

@claude-flow/browser (v3 monorepo, 59 MCP tools) ruflo-browser plugin (23 MCP tools)
agent-browser dependency locked to ^0.6.0, upstream was at 0.27.0 (21 versions behind) at ADR write time shells out to whatever's on PATH
What's actually running in this project not this one confirmed this one, by counting the live registered mcp__ruflo__browser_* tools (exactly 23)

ADR-122's own Phase 0 (its "highest-yield, lowest-risk" first step) is specifically:
bump the locked agent-browser version, converge the two systems, and — this is the
relevant part — "ruflo doctor reports agent-browser version and warns when below
0.27."
That's the same health-reporting capability being asked for here, just scoped to
land inside ruflo itself instead of ak.

So there's a real design that already intends to close this gap — it's just stalled
(Proposed since May, not shipped by September). Flagging this so the decision is informed:
ak including agent-browser/vibium now is still valuable regardless of whether/when
ADR-122 ships (health-checking a third-party binary a project depends on is ak's job
either way) — but it's bridging a known gap rather than inventing new scope from nothing.

Three more gaps, found the same session, all verified against your source

Not padding — these came up solving the actual problem above and are all confirmed
against the installed 4.0.0-alpha.44 source, not guessed. Splitting them out so they
don't dilute the main ask, but flagging them since you asked for detail and they're the
same shape of "ak already has 90% of this, just missing a piece."

1. ak x mcp pick's legacy-key migration is blind to project-scoped registrations

src/lib/mcp.mjs's registrationStatus() / register() only read/write
~/.claude.json's top-level mcpServers (claudeUserMcpPath() →
path.join(home, '.claude.json'), global scope). This project's actual live MCP
registration — the one every mcp__ruflo__browser_* call in this whole investigation
went through — lives at ~/.claude.json → projects["/workspaces/turbo-flow-wsl"] .mcpServers.ruflo, plus a project-local .mcp.json. Neither is visible to
registrationStatus()'s legacyRuflo check or x mcp pick's migration.

Concretely: this machine's global scope is already clean (claude-flow registered,
no legacy ruflo key) — but this project's own registration is still on the legacy
ruflo key, entirely invisible to the tooling that's supposed to catch that. Confirmed
by reading ~/.claude.json directly: the project-scoped entry existed with env: {},
silently shadowing every env var the project's own .mcp.json declared — which is what
actually broke the fix for the Chrome-binary path above and needed manual patching
across both files.

2. ak x daemon-gc doesn't cover orphaned ruflo mcp start processes

x/daemon-gc.mjs + lib/daemons.mjs already do exactly the right thing — TTL-based
staleness, workspace-gone detection, list-by-default/--kill-to-reap — for ak's own
background daemons. Zero references to mcp start process discovery anywhere in
daemons.mjs. Also worth a look while you're in there: a separate bead from this same
project's internal tracker found a stray ruflo daemon start --foreground --workspace <unrelated-nested-path> process, and daemon-gc's own "1-per-project is normal"
heuristic considered it healthy and left it running — it doesn't currently catch a
daemon that's alive and within its TTL, but pointed at the wrong project. Meanwhile, watched /mcp reconnect (in Claude Code, not ak) spawn a new
ruflo mcp start process pair on every reconnect without killing the one it replaced —
confirmed live: 3 separate spawns in one terminal within 6 minutes, all still running
(kill -0 succeeded on each) after later ones had superseded them. Already reported the
underlying leak as Claude Code product feedback (not ak's or ruflo's bug), but
daemon-gc's existing pattern would be the natural place to also let people clean up the
resulting mess, the same way it already does for ak's own daemons.

3. ak's own ALLOW_SCRIPTS allowlist is missing @anthropic-ai/claude-code

This one's the most concrete — it's the literal cause of a "claude native binary not
installed" failure hit and hand-fixed this same session. lib/heal.mjs has:

// Packages whose install scripts must run for natives to build (npm >=11.17
// blocks them by default). Curated on the live 3.28/3.12.2 upgrade.
const ALLOW_SCRIPTS = [
  'ruflo', 'agentic-qe', '@claude-flow/cli', 'better-sqlite3', 'hnswlib-node',
  'agentdb', 'agentic-flow', 'argon2', 'onnxruntime-node', 'sharp', 'protobufjs',
  '@google/genai', 'tldjs', 'vibium',
].join(',');

@anthropic-ai/claude-code isn't in it. Its postinstall (install.cjs) is exactly the
kind of native-binary-fetching script npm ≥11.17 blocks by default — same failure mode
ALLOW_SCRIPTS already exists to prevent for 14 other packages. Hit this directly: an
auto-update bumped the wrapper package but the postinstall that fetches/links the
matching native binary got silently blocked, and dsp/claude broke with "native binary
not installed" until fixed by hand (adding @anthropic-ai/claude-code to
~/.npmrc's allow-scripts). Since ak is what most people would reach for to heal a
broken Claude Code CLI install, and it already has the exact mechanism to prevent this
one specific package from ever hitting it — seems like a one-line fix with real, verified
impact.

This isn't a one-off, either — this project's own internal tracker shows the same class
of break recurring over weeks. One earlier incident hit and fixed the exact same
mechanism as this session — a fresh npm install -g blocked by the allow-scripts policy
— with the identical fix, about two weeks before this session independently
rediscovered it from scratch. That's the confirmed, reproducible failure mode
ALLOW_SCRIPTS would directly prevent.

A separate, earlier incident found something related but not fully
explained: 3 breaks in one session where the claude.exe wrapper got reset to a bare
failure stub while the actual ~330MB native binary stayed completely intact — not
obviously a full reinstall each time. Its author ruled out cron, systemd, shell hooks,
.bashrc/aliases, and devcontainer lifecycle scripts, killed a stray unrelated daemon as
a leading suspect (didn't help), and gave up identifying the real trigger, landing on a
.bashrc self-heal workaround instead — so it's honest to say ALLOW_SCRIPTS is
confirmed to fix the ltpp/this-session pattern, not necessarily gdrr's deeper mystery.
Still worth including as evidence this general area (Claude Code's native binary going
missing) has cost real time here more than once.

References

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions