Skip to content

fix(natives): resolve better-sqlite3 by fresh fs walk, not cached resolver - #25

Merged
pacphi merged 1 commit into
mainfrom
fix/natives-stale-resolver-cache
Jul 16, 2026
Merged

pacphi merged 1 commit into
mainfrom
fix/natives-stale-resolver-cache

Conversation

@pacphi

@pacphi pacphi commented Jul 16, 2026

Copy link
Copy Markdown
Owner

Problem

ak sync could end with a false failure even when everything was actually healthy:

✓ natives: …/ruflo/node_modules/agentdb: native installed; …/agentic-flow/node_modules/agentdb: native built…; agentic-qe: native installed
…
✗ still failing: [natives] 1/1 agentdb location(s) on WASM fallback (data-loss writes)

The heal step reported the location native, yet the final convergence proof reported WASM fallback — and a fresh node process confirmed the on-disk state was native and healthy. So the end state was fine; the proof lied.

Root cause

bsq3Root() resolved better-sqlite3 via createRequire().resolve(). Node's resolution cache (Module._pathCache / _realpathCache) is process-wide and goes stale after an in-process npm install reshapes node_modules. During one sync:

  1. healNatives() resolves + builds native better-sqlite3 (priming the cache; the tree has a nested agentic-flow/node_modules/agentdb).
  2. healAidefence() runs npm install @claude-flow/aidefence into the ruflo tree → dedupes the nested agentdb away.
  3. The convergence proof calls bsq3Root() again → gets the cached root pointing at the now-deleted nested path → binding check fails → false WASM.

agentdbLocations() uses plain fs.existsSync and correctly saw the tree shrink 2→1; only the resolver-backed lookup was stale.

Fix

bsq3Root() now walks up the node_modules chain with fs.existsSync on every call — the Node-resolution equivalent, but reading disk fresh, so it's immune to in-process tree mutation (matching how agentdbLocations() already reads).

  • Identical behavior for the resolvable / not-resolvable / native / WASM cases (existing tests unchanged).
  • Adds coverage for hoisted resolution (the real ruflo/node_modules/agentdb → ruflo/node_modules/better-sqlite3 layout) and the nested→hoisted mid-process move that triggered the bug.

Verification

npm run check green (exit 0): typecheck, eslint, markdownlint, build, 76 kit + 20 statusline tests. Verified live on the reporting machine: ak status shows ✓ natives native better-sqlite3 in 1 agentdb location(s) and ak sync now reports ✓ converged — no failing subsystems.

🤖 Generated with Claude Code

…olver

`ak sync` intermittently ended with a false "[natives] 1/1 agentdb
location(s) on WASM fallback (data-loss writes)" even though the heal step
had just reported the location native — and a fresh process confirmed it
native. Root cause: bsq3Root() used createRequire().resolve(), whose
process-wide resolution cache (Module._pathCache/_realpathCache) goes stale
after an in-process `npm install` (the aidefence heal / a version upgrade)
dedupes the tree. The cached root pointed at a nested better-sqlite3 that
npm had just removed, so the binding check on that dead path failed —
a false negative in sync's convergence proof. agentdbLocations() reads with
plain fs and was already correct; only the resolver-backed lookup lied.

Fix: bsq3Root() now walks up the node_modules chain with fs.existsSync on
every call — the node-resolution equivalent, but reading disk fresh, immune
to in-process tree mutation. Same semantics for the resolvable/native/WASM
cases (existing tests unchanged); adds coverage for hoisted resolution and
the nested→hoisted mid-process move that triggered the bug.
@pacphi
pacphi merged commit 798f43c into main Jul 16, 2026
11 checks passed
@pacphi
pacphi deleted the fix/natives-stale-resolver-cache branch July 16, 2026 06:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant