Skip to content

feat(hooks): two-tier hook evaluator — a regex floor plus a semantic tier, off by default #2215

feat(hooks): two-tier hook evaluator — a regex floor plus a semantic tier, off by default

feat(hooks): two-tier hook evaluator — a regex floor plus a semantic tier, off by default #2215

Workflow file for this run

name: CI
on:
push:
branches: [main]
# `main` AND the long-lived branches other PRs stack onto. A pull request is
# tested against the branch it will actually merge into, and a stacked PR
# targeting anything but `main` matched nothing here — so #730, thirty commits
# of SDK work, ran no unit tests, no build and no lint at all. The only signal
# it produced was the daemon cross-compile, and only because it touched
# `crates/`. A PR that cannot go red is not a reviewed PR.
pull_request:
branches: [main, feat/fp-cli]
concurrency:
group: ci-${{ github.ref }}
cancel-in-progress: true
jobs:
quality:
runs-on: ubuntu-latest
timeout-minutes: 10
env:
FAILPROOFAI_TELEMETRY_DISABLED: "1"
steps:
- uses: actions/checkout@v7.0.1
- uses: oven-sh/setup-bun@v2
with:
bun-version: latest
# Restore-only here and in every other job; `quality` alone writes, below.
# All six jobs used to share one read-write `actions/cache@v6` on this key,
# so all six raced to upload the same 401 MB entry and five of them lost
# the race — paying the upload to be told the key already existed. One
# writer is all a shared key can use.
- uses: actions/cache/restore@v6
id: bun-cache
with:
path: ~/.bun/install/cache
key: bun-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-${{ runner.os }}-
# `--ignore-scripts`, here and in every other job, because package.json's
# `prepare` is `bun run build` — so a plain `bun install` runs a full
# Next.js production build (~28s: 14s compile + 13s TypeScript) as an
# install lifecycle hook, in six of this workflow's eight jobs. Nothing in
# this job reads `dist/`: tsconfig.json excludes it, eslint.config.mjs
# ignores it, and no app/ file references Next's generated route types.
# rust-quality has guarded this way since it landed, and its install takes
# one second against the thirty-three this one used to.
- name: Install dependencies
uses: nick-fields/retry@v4
with:
max_attempts: 3
timeout_minutes: 5
command: bun install --frozen-lockfile --ignore-scripts
- uses: actions/cache/save@v6
if: steps.bun-cache.outputs.cache-hit != 'true'
with:
path: ~/.bun/install/cache
key: bun-${{ runner.os }}-${{ hashFiles('bun.lock') }}
- name: Check version consistency
run: |
ROOT_VERSION=$(jq -r .version package.json)
echo "Root version: $ROOT_VERSION"
MISMATCH=0
for pkg in packages/*/package.json; do
[ -f "$pkg" ] || continue
PKG_VERSION=$(jq -r .version "$pkg")
if [ "$PKG_VERSION" != "$ROOT_VERSION" ]; then
echo "::error file=$pkg::Version mismatch: $pkg has $PKG_VERSION, expected $ROOT_VERSION"
MISMATCH=1
fi
done
# Check optionalDependencies in wrapper
for dep_version in $(jq -r '.optionalDependencies // {} | values[]' packages/wrapper/package.json 2>/dev/null || true); do
if [ "$dep_version" != "$ROOT_VERSION" ]; then
echo "::error file=packages/wrapper/package.json::Dependency version mismatch: $dep_version, expected $ROOT_VERSION"
MISMATCH=1
fi
done
# The daemon binaries DO ship as npm platform packages
# (@failproofai/failproofaid-<os>-<arch>), but their pins are injected
# into package.json at publish time by
# scripts/build-daemon-packages.mjs — the same invocation that
# publishes them, so they cannot drift — and are deliberately absent
# from the committed tree. Nothing to check here.
#
# The Cargo version still has to match, because the release tag the
# CLI builds its download URL from is the npm version, and the binary
# at that URL reports the Cargo one.
# Check the Cargo workspace version (failproofaid) against root package.json
if [ -f Cargo.toml ]; then
CARGO_VERSION=$(grep -m1 '^version = ' Cargo.toml | sed -E 's/version = "(.*)"/\1/')
if [ "$CARGO_VERSION" != "$ROOT_VERSION" ]; then
echo "::error file=Cargo.toml::Version mismatch: Cargo.toml has $CARGO_VERSION, expected $ROOT_VERSION"
MISMATCH=1
fi
fi
if [ "$MISMATCH" -eq 1 ]; then
echo "::error::Version mismatch detected across package.json files"
exit 1
fi
echo "All versions match: $ROOT_VERSION"
- name: Lint
uses: nick-fields/retry@v4
with:
max_attempts: 3
timeout_minutes: 5
command: bun run lint
- name: Type check
uses: nick-fields/retry@v4
with:
max_attempts: 3
timeout_minutes: 5
command: bunx tsc --noEmit
rust-quality:
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v7.0.1
with:
# This job runs `cargo clippy`/`cargo test` over the full
# dependency tree, executing third-party build scripts. The
# default (`true`) would leave GITHUB_TOKEN in .git/config where
# any of them could read it; nothing here needs push access.
persist-credentials: false
# On `pull_request`, actions/checkout builds the merge commit, whose
# first parent is the base — so depth 2 is exactly enough for the
# diff below, without fetching the branch's history.
fetch-depth: 2
# Two gates in one, both answering "is there Rust work to do here".
#
# The first is the original: stage 1 landed an empty Cargo workspace so
# the CI plumbing could go green before any Rust existed, and
# `cargo build/clippy/test --workspace` (and even `cargo fmt --all`) all
# hard-error on a zero-member workspace rather than no-opping cleanly.
# All three crates exist now, so on its own this always answers true.
#
# The second is why it is still here. This is the longest job in CI, so it
# sets the wall clock for EVERY pull request — including the many that
# touch no Rust at all, where it restores a cache and recompiles a
# 231-crate dependency tree to check nothing. Gating on the diff leaves
# the job reporting a status (no `needs:` edge, no new job, nothing
# serialised behind it) while finishing in seconds on a TypeScript-only
# branch. `push` is never gated: main always gets the full check, which is
# also what keeps the cache warm for everyone else.
- name: Detect Rust changes
id: crates
run: |
if ! ls crates/*/Cargo.toml >/dev/null 2>&1; then
echo "present=false" >> "$GITHUB_OUTPUT"
echo "No crates/*/Cargo.toml on this ref — rust-quality has nothing to check."
exit 0
fi
if [ "${{ github.event_name }}" != "pull_request" ]; then
echo "present=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# Anything that can change what clippy or the tests see — which is
# NOT only Rust. `cargo test --workspace` spawns the real TS worker
# (`bun bin/failproofai-worker.mjs`, from server.rs's live end-to-end
# test), and that worker runs raw TypeScript, resolving src/hooks'
# real dependency tree at runtime rather than a bundle's. So a change
# under src/hooks, to the worker entrypoint, or to the dependency
# graph it resolves against can break this job with no Rust involved
# at all — the failure surfaces as "worker process exited before
# creating its socket", which is exactly the class a gate that only
# matched Rust paths would have let through silently.
# ci.yml is in the list so a change to this gate always exercises it.
if git diff --name-only HEAD^1 HEAD | grep -qE '^(crates/|src/hooks/|bin/failproofai-worker\.mjs$|package\.json$|bun\.lock$|Cargo\.toml$|Cargo\.lock$|rust-toolchain\.toml$|\.github/workflows/ci\.yml$)'; then
echo "present=true" >> "$GITHUB_OUTPUT"
else
echo "present=false" >> "$GITHUB_OUTPUT"
echo "No Rust-affecting paths in this PR — skipping fmt/clippy/test."
fi
- if: steps.crates.outputs.present == 'true'
run: rustup show
# cargo test spawns the real TS worker via `bun bin/failproofai-worker.mjs`
# (crates/failproofaid/src/server.rs's live end-to-end test) — bun has to
# be on PATH for that test, not just for the TS-side jobs.
- if: steps.crates.outputs.present == 'true'
uses: oven-sh/setup-bun@v2
with:
bun-version: latest
# …and bun on PATH is not enough on its own. That worker runs the raw
# TypeScript, so it resolves the handler's real dependency tree at
# runtime rather than a bundle's. The moment anything under src/hooks
# imports a third-party package, this job fails with a bun ENOENT that
# reads like a Rust problem — the failure surfaces as "worker process
# exited before creating its socket". Production is unaffected, because
# dist/worker.mjs bundles those deps; only this path needs them on disk.
- if: steps.crates.outputs.present == 'true'
uses: actions/cache/restore@v6
with:
path: ~/.bun/install/cache
key: bun-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-${{ runner.os }}-
- if: steps.crates.outputs.present == 'true'
name: Install worker dependencies
run: bun install --frozen-lockfile --ignore-scripts
# Restore on every run, SAVE ONLY ON MAIN — the split `build-daemon.yml`
# already uses, for a second reason that turned out to matter more.
#
# A combined `actions/cache@v6` writes a ref-scoped copy from every branch
# that misses the exact key, and this entry carries `target/`, so each copy
# is 1.5-2.3 GiB. Five PR refs held one at once (677, 679, 680, 681 and
# main) — ~10.7 GiB of a repo cache that GitHub caps at 10 GiB, which puts
# the store permanently in LRU eviction.
#
# What that evicted was not another cargo build. It was the 13 KB
# translation cache, touched once every 24 hours by the nightly
# `translate-docs` run and therefore always the least-recently-used thing
# in the store. Losing it re-translated all 48 pages into all 14 languages
# the next morning: ~125 runner-minutes and a full LLM pass per language,
# against a 4-minute baseline when the cache survives. Six consecutive days
# of it, Aug 6-11, cost ~750 runner-minutes and six full translation passes
# through the gateway.
#
# Restoring without saving costs a PR whose `Cargo.lock` moved a rebuild
# from a stale-but-close main cache — which is already what `restore-keys`
# hands it today.
#
# That fixed the MULTIPLICATION and left the SIZE, and the size was the
# bigger half. One entry reached 5,727 MB — 57% of the whole 10 GiB quota
# in a single key, so the store stayed in permanent LRU eviction against
# the translation cache anyway — and restoring it took 127 SECONDS, more
# wall clock than the 74s `cargo test` it was there to avoid. A cache that
# costs more to move than the work it replaces is not a cache.
#
# The cause is `path: target` taken literally: it archives every
# intermediate the workspace has ever produced, including this workspace's
# own crates, which recompile in seconds and are the artifacts most likely
# to be stale. rust-cache caches the dependency artifacts and prunes the
# rest. It also handles the save gate directly, so the paired
# `cache/save` step below this is gone.
- if: steps.crates.outputs.present == 'true'
uses: Swatinem/rust-cache@v2
with:
save-if: ${{ github.event_name == 'push' && github.ref == 'refs/heads/main' }}
- name: cargo fmt --check
if: steps.crates.outputs.present == 'true'
run: cargo fmt --all -- --check
- name: cargo clippy
if: steps.crates.outputs.present == 'true'
run: cargo clippy --workspace --all-targets -- -D warnings
- name: cargo test
if: steps.crates.outputs.present == 'true'
run: cargo test --workspace
# The `fp` CLI (PyPI: fp-cloud-cli). The only Python in this repo, and the only job that
# tests it — none of the bun/cargo jobs above look at fp-cloud-cli/ at all. It is matrixed
# across the Python versions pyproject.toml's requires-python advertises, because
# claiming >=3.10 and testing only one of them is how a 3.10 user finds the break.
#
# That comment was true of the SDK's matrix and not of this one, which ran 3.10
# and 3.13 only while the classifiers claimed 3.10–3.13 — so a resolver picking
# a different transitive set on 3.11 or 3.12 broke those users with fully green
# CI and a wheel whose own metadata said they were supported. All four now, and
# `__tests__/ci/fp-cloud-cli-workflows.test.ts` pins the list against the classifiers
# so the two cannot drift apart again.
fp-cloud-cli:
runs-on: ubuntu-latest
# a uv sync plus pytest across four interpreters; the bound exists so a stalled
# package mirror cannot hold a release for six hours (#726).
timeout-minutes: 15
defaults:
run:
working-directory: fp-cloud-cli
strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.11", "3.12", "3.13"]
steps:
- uses: actions/checkout@v7.0.1
- uses: astral-sh/setup-uv@v7
with:
enable-cache: true
cache-dependency-glob: fp-cloud-cli/uv.lock
- name: Install dependencies
run: uv sync --locked --extra dev --python ${{ matrix.python-version }}
# FP_CLI_REQUIRE_CONTRACT makes test_fp_home_contract.py FAIL rather than skip
# when it cannot read src/hooks/fp-home.ts. That file is the register for a
# home this CLI writes a credential into, and the assertions against it are
# the only thing keeping the two sides from drifting — before they existed, a
# rename on either side left both suites green. A guard that can quietly
# degrade to a skip is the same guard the SDK's spool contract used to be.
- name: Test
env:
FP_CLI_REQUIRE_CONTRACT: "1"
run: uv run pytest tests/ -q
- name: Build the distribution
run: uv build
# A stale `include = [...]` in [tool.setuptools.packages.find] builds a
# SUCCESSFUL but EMPTY wheel — pip installs it and the console script then
# ImportErrors. Nothing else in this job would notice, so assert the payload.
- name: Verify the wheel actually contains the package
run: |
python3 - <<'PY'
import glob, sys, zipfile
wheel = glob.glob("dist/*.whl")[0]
names = zipfile.ZipFile(wheel).namelist()
modules = [n for n in names if n.startswith("fp_cli/") and n.endswith(".py")]
print(f"{wheel}: {len(modules)} modules")
if len(modules) < 20:
sys.exit(f"wheel looks empty — only {len(modules)} modules under fp_cli/")
PY
# `fp` is the contract users type. Prove the entry point resolves from a clean
# install of the built artifact, not from the source tree.
- name: Smoke-test the installed console script
run: |
uv venv /tmp/fp-smoke
VIRTUAL_ENV=/tmp/fp-smoke uv pip install dist/*.whl
/tmp/fp-smoke/bin/fp --version
/tmp/fp-smoke/bin/fp help > /tmp/fp-help.txt
if grep -qi agenteye /tmp/fp-help.txt; then
echo '::error::retired product name present in the fp help output'
exit 1
fi
# The telemetry SDK (PyPI: failproofai-sdk, import failproofai_sdk). Second Python
# component, sibling of fp-cloud-cli above, and deliberately its own job: it shares no
# lockfile, no dependencies and no release cadence with the CLI.
#
# The matrix is wider than fp-cloud-cli's two versions on purpose. This package declares
# no dependencies at all, so it has no third-party floor quietly constraining which
# interpreters it is really exercised on — and it is installed into other people's
# agent processes, whose Python version we do not choose. Every version
# requires-python advertises is tested.
failproofai-sdk:
runs-on: ubuntu-latest
# the same, across five — every version requires-python advertises; the bound exists so a stalled
# package mirror cannot hold a release for six hours (#726).
timeout-minutes: 10
defaults:
run:
working-directory: sdk/python
strategy:
fail-fast: false
matrix:
python-version: ["3.10", "3.11", "3.12", "3.13", "3.14"]
steps:
- uses: actions/checkout@v7.0.1
- uses: astral-sh/setup-uv@v7
with:
enable-cache: true
cache-dependency-glob: sdk/python/uv.lock
- name: Install dependencies
run: uv sync --locked --extra dev --python ${{ matrix.python-version }}
# FAILPROOFAI_SDK_REQUIRE_CONTRACT makes test_spool_contract.py FAIL rather than
# skip when it cannot find crates/fpai-collect or src/hooks/fp-home.ts. Those
# assertions are the only check that this SDK writes where the daemon reads, and
# a guard that can silently degrade to a skip is not a guard — its predecessor in
# the agenteye repo skipped in every CI run for exactly this reason.
- name: Test
env:
FAILPROOFAI_SDK_REQUIRE_CONTRACT: "1"
run: uv run pytest tests/ -q
- name: Build the distribution
run: uv build
# A stale `include = [...]` in [tool.setuptools.packages.find] builds a
# SUCCESSFUL but EMPTY wheel — pip installs it and the import then fails.
# py.typed goes the same way if package-data stops naming it, and nothing
# else in this job would notice either. Assert the payload.
- name: Verify the wheel actually contains the package
run: |
python3 - <<'PY'
import glob, sys, zipfile
wheel = glob.glob("dist/*.whl")[0]
names = zipfile.ZipFile(wheel).namelist()
modules = [n for n in names if n.startswith("failproofai_sdk/") and n.endswith(".py")]
print(f"{wheel}: {len(modules)} modules")
if len(modules) < 7:
sys.exit(f"wheel looks empty — only {len(modules)} modules under failproofai_sdk/")
if not any(n.endswith("py.typed") for n in names):
sys.exit("py.typed is missing from the wheel — the type hints do not count without it")
if not any(n.endswith("LICENSE") for n in names):
sys.exit("LICENSE is missing from the wheel")
PY
# `--no-deps` is the assertion, not an optimisation. "Zero dependencies" is the
# reason this package is safe to drop into someone else's agent, so prove it
# against the built artifact rather than the source tree: install it with nothing
# else present, emit real events, and read them back off disk.
- name: Smoke-test the installed package with no dependencies
run: |
uv venv /tmp/sdk-smoke
VIRTUAL_ENV=/tmp/sdk-smoke uv pip install --no-deps dist/*.whl
FAILPROOFAI_HOME=/tmp/sdk-spool /tmp/sdk-smoke/bin/python -c "
import failproofai_sdk as s
print(s.__version__)
s.event.agent_start(session_id='smoke', agent_id='a', goal='ci')
s.event.tool_use(session_id='smoke', agent_id='a', tool_name='t', tool_call_id='c')
s.event.tool_result(session_id='smoke', agent_id='a', tool_name='t', tool_call_id='c', output='ok')
"
python3 - <<'PY'
import glob, json, sys
# $FAILPROOFAI_HOME/custom-agents/events — the `custom-agents` segment
# is appended unconditionally by the resolver, which is what makes the
# spool impossible to place outside the umbrella.
batches = glob.glob("/tmp/sdk-spool/custom-agents/events/*.jsonl")
if not batches:
sys.exit("the installed wheel wrote no event batch")
events = [json.loads(line) for p in batches for line in open(p) if line.strip()]
types = {e["type"] for e in events}
if types != {"agent_start", "tool_use", "tool_result"}:
sys.exit(f"unexpected events from the installed wheel: {sorted(types)}")
# Emitted by the artifact, so this also proves the wire format survived
# packaging rather than only surviving an in-tree import.
if not all("environment" in e for e in events):
sys.exit("an event reached disk without the `environment` field")
print(f"{len(events)} events written by the installed artifact")
PY
# The adapters' ONLY automated evidence. It is a separate job, not a leg of the
# matrix above, because installing five agent frameworks costs minutes and gains
# nothing from being repeated across five interpreters — the adapters bind to
# framework APIs, not to interpreter version.
#
# WHY THIS EXISTS: tests/integrations/* already honour
# AGENTEYE_TESTS_REQUIRE_FRAMEWORKS, and their comments say "CI leg sets
# AGENTEYE_TESTS_REQUIRE_FRAMEWORKS=1". No such leg was ever added, and the
# frameworks live in extras that `--extra dev` does not pull, so all four
# modules skipped at import in every run: 168 test functions across 6,071 lines,
# green, never executed. An adapter could break against a new framework release
# and nothing here would say so. That is the same "a guard that can silently
# degrade to a skip is not a guard" this file already hardened
# test_spool_contract.py against — the integration suites were simply missed.
failproofai-sdk-integrations:
runs-on: ubuntu-latest
timeout-minutes: 20
defaults:
run:
working-directory: sdk/python
steps:
- uses: actions/checkout@v7.0.1
- uses: astral-sh/setup-uv@v7
with:
enable-cache: true
cache-dependency-glob: sdk/python/uv.lock
# Every framework extra. `--locked` keeps this honest: the lockfile already
# resolves all five, so a drift here fails rather than silently re-resolving.
- name: Install dependencies (all framework extras)
run: >-
uv sync --locked --extra dev
--extra langchain --extra langgraph --extra crewai
--extra llamaindex --extra pydantic-ai
--python 3.12
# AGENTEYE_TESTS_REQUIRE_FRAMEWORKS turns each module's import skip into a
# hard failure. Without it a botched install above would read as "4 skipped"
# and the job would pass having tested nothing — which is exactly the state
# this job was added to end.
- name: Test the framework adapters
env:
AGENTEYE_TESTS_REQUIRE_FRAMEWORKS: "1"
run: uv run pytest tests/integrations -q
# The TypeScript telemetry SDK (`@failproofai/sdk`). Separate from `quality`
# and `test` because it is its own npm package with its own lockfile, its own
# tsconfig and its own vitest config — running it inside the root project's
# jobs would mean the root's dependency tree decided whether this package's
# zero-dependency claim holds.
#
# The matrix is Node's supported majors, not a single version. This package's
# floor is 20.9 and its `using` support, `AsyncLocalStorage.enterWith`,
# `worker_threads` resource limits and `Symbol.dispose` shim all behave
# differently across that range — which is precisely the range a customer's
# agent runs on.
failproofai-ts-sdk:
runs-on: ubuntu-latest
timeout-minutes: 15
defaults:
run:
working-directory: sdk/typescript
strategy:
fail-fast: false
matrix:
include:
# The `engines` FLOOR. This leg proves the published artifact runs on
# it, and deliberately does not run the suite: vitest 5 pulls vite 8,
# which pulls rolldown, which imports `styleText` from `node:util` —
# added in Node 20.12. The TEST RUNNER's floor is not the PACKAGE's
# floor, and the honest way to say so is to keep the floor at 20.9 and
# prove it with the thing a consumer actually gets.
#
# Pinning the runner back to something 20.9 can load is the tail
# wagging the dog: it means carrying the CVEs vitest 5 fixed (one
# Critical, one High) so that a test runner can start on a Node
# release nobody runs the tests on.
- node-version: "20.9"
suite: false
- node-version: "20.x"
suite: true
- node-version: "22.x"
suite: true
- node-version: "24.x"
suite: true
steps:
- uses: actions/checkout@v7.0.1
- uses: actions/setup-node@v5
with:
node-version: ${{ matrix.node-version }}
cache: npm
cache-dependency-path: sdk/typescript/package-lock.json
- name: Install dependencies
uses: nick-fields/retry@v4
with:
max_attempts: 3
timeout_minutes: 5
command: cd sdk/typescript && npm ci --no-audit --no-fund
- name: Typecheck
if: matrix.suite
run: npm run typecheck
- name: Lint
if: matrix.suite
run: npm run lint
# Runs on EVERY leg, including the floor: `tsc` is the thing that produces
# what ships, so "it builds on 20.9" is a claim worth checking there.
- name: Build
run: npm run build
# The sandbox suite needs `dist/` present — the evaluator sandbox is a
# real `worker_threads` entry and cannot load a `.ts` file — and it
# asserts rather than skips when it is missing, so the build above is a
# prerequisite rather than a duplicate.
- name: Test
if: matrix.suite
run: npx vitest run
# Everything above ran against the source tree. These two steps run
# against the ARTIFACT, because the failures they catch — a missing export
# condition, a CommonJS build Node reads as ESM, a `dist/` path the files
# list does not ship — are invisible from inside the package and total
# from outside it.
- name: Pack
run: npm pack --pack-destination /tmp
- name: Smoke-test the packed tarball with no dependencies
run: |
mkdir -p /tmp/ts-sdk-smoke && cd /tmp/ts-sdk-smoke
npm init -y >/dev/null
# `--omit=optional --omit=peer` is the assertion, not an optimisation.
# "Zero dependencies" is the reason this package is safe to drop into
# someone else's agent, so prove it against the built artifact: install
# it with nothing else present, emit real events, and read them back
# off disk.
npm install --no-audit --no-fund --omit=optional --omit=peer /tmp/failproofai-sdk-*.tgz
test ! -d node_modules/@failproofai/sdk/node_modules \
|| { echo "the published package brought transitive dependencies"; exit 1; }
cat > esm.mjs <<'EOF'
import * as fp from "@failproofai/sdk";
fp.configure({ baseDir: process.env.SPOOL });
await fp.agent("smoke", { goal: "ci" }, async () => {
await fp.toolCall("t", { input: { q: 1 } }, () => "ok");
});
await fp.flush();
console.log(fp.version);
EOF
SPOOL=/tmp/ts-sdk-spool node esm.mjs
cat > cjs.cjs <<'EOF'
const fp = require("@failproofai/sdk");
fp.configure({ baseDir: process.env.SPOOL });
fp.event.agentStart({ sessionId: "cjs", goal: "ci" });
fp.flushSync();
console.log(fp.version);
EOF
SPOOL=/tmp/ts-sdk-spool node cjs.cjs
# The evaluator loads from its own subpath, and its CLI has to be
# executable — a `bin` that is not fails on every platform where the
# installer links rather than copies.
node -e "const e = require('@failproofai/sdk/evaluator'); if (typeof e.Evaluator !== 'function') throw new Error('evaluator subpath is broken')"
npx --no-install failproofai-evaluator --help > /dev/null
node - <<'EOF'
const { readdirSync, readFileSync } = require("node:fs");
const { join } = require("node:path");
const dir = "/tmp/ts-sdk-spool/events";
const events = readdirSync(dir)
.filter((f) => f.endsWith(".jsonl"))
.flatMap((f) => readFileSync(join(dir, f), "utf8").split("\n").filter(Boolean))
.map((line) => JSON.parse(line));
const types = new Set(events.map((e) => e.type));
const want = ["agent_start", "agent_end", "tool_use", "tool_result"];
for (const type of want) {
if (!types.has(type)) throw new Error(`the installed artifact never wrote ${type}`);
}
// Emitted by the artifact, so this also proves the wire format
// survived packaging rather than only surviving an in-tree import.
if (!events.every((e) => "environment" in e && "session_id" in e)) {
throw new Error("an event reached disk missing a required field");
}
console.log(`${events.length} events written by the installed artifact`);
EOF
# The evaluator sandbox in an INSTALLED package resolves its worker
# through the package's own `./sandbox-worker` export, with no env
# override in sight. That resolution is the one part of the sandbox that
# cannot be exercised from inside this repository, and a failure in it
# means managed evaluations refuse to run for every customer.
- name: Verify the evaluator sandbox resolves from an installed package
run: |
cd /tmp/ts-sdk-smoke
node - <<'EOF'
const { compileEvaluator, sessionTranscriptFromWire } = require("@failproofai/sdk/evaluator");
const session = sessionTranscriptFromWire({
schema_version: "2",
assignment_id: "a", session_id: "s", session_revision_id: "r",
agent_id: "main", environment: "dev",
started_at: "2026-01-01T00:00:00.000000Z",
ended_at: "2026-01-01T00:01:00.000000Z",
event_count: 1,
events: [{ id: "1", ts: "2026-01-01T00:00:01.000000Z", event_type: "tool_use", payload: {} }],
});
compileEvaluator("EvalResult({ score: Score(session.count('tool_use') > 0 ? 1 : 0) })", { evalKey: "k" })(session)
.then((result) => {
if (result.score.value !== 1) throw new Error(`unexpected score ${result.score.value}`);
console.log("sandbox resolved and evaluated from the installed package");
})
.catch((error) => { console.error(error); process.exit(1); });
EOF
# The TypeScript SDK's framework adapters, against REAL framework releases —
# the counterpart of `failproofai-sdk-integrations` above. `failproofai-ts-sdk`
# runs the adapters with no framework installed, which proves their logic and
# nothing about whether it ever reaches a framework: the first release passed
# 242 unit tests with every adapter recording nothing in an ES-module app.
#
# Each fixture under `sdk/typescript/integration/fixtures/` is a consumer
# project with its own lockfile pinning one framework release. The packed
# tarball is extracted into each, and one agent is run as BOTH an ES module and
# CommonJS, because the two module systems load different copies of a
# dual-published framework. A fixture that fails to install fails the job —
# there is no skip path to read as green.
failproofai-ts-sdk-integrations:
name: failproofai-ts-sdk-integrations (${{ matrix.shard }}, node ${{ matrix.node-version }})
runs-on: ubuntu-latest
timeout-minutes: 30
defaults:
run:
working-directory: sdk/typescript
strategy:
fail-fast: false
matrix:
# Sharded by what a shard needs installed, because the whole suite —
# four frameworks at both ends of their ranges, Bun and Deno parity over
# every fixture, and five real `next build`s — is ~23 CPU-minutes, too
# much for one runner inside a timeout. Each shard installs only its own
# fixtures (FAILPROOFAI_IT_FIXTURES) and runs only its own files.
#
# frameworks: Node's oldest and newest supported majors — the ESM/CJS
# split this job exists for behaves differently once `require(esm)` is
# unflagged. runtimes and nextjs: one Node, because what they vary is
# the runtime or the bundler, not Node.
include:
- shard: frameworks
node-version: "20.x"
fixtures: ai-4,ai-5,ai-6,ai-7,langchain-0.3,langchain-1,langchain-dup-core,mastra-0,mastra-1,llamaindex-0.11,llamaindex-0.12,types,vanilla
files: integration/ai.test.ts integration/langchain.test.ts integration/mastra.test.ts integration/mastra-coverage.test.ts integration/llamaindex.test.ts integration/types.test.ts integration/vanilla.test.ts
- shard: frameworks
node-version: "24.x"
fixtures: ai-4,ai-5,ai-6,ai-7,langchain-0.3,langchain-1,langchain-dup-core,mastra-0,mastra-1,llamaindex-0.11,llamaindex-0.12,types,vanilla
files: integration/ai.test.ts integration/langchain.test.ts integration/mastra.test.ts integration/mastra-coverage.test.ts integration/llamaindex.test.ts integration/types.test.ts integration/vanilla.test.ts
- shard: runtimes
node-version: "24.x"
fixtures: ai-4,ai-5,ai-6,ai-7,langchain-0.3,langchain-1,mastra-0,mastra-1,llamaindex-0.11,llamaindex-0.12,runtimes
files: integration/runtimes.core.test.ts integration/runtimes.bun.test.ts integration/runtimes.deno.test.ts
- shard: nextjs
node-version: "24.x"
fixtures: nextjs,langchain-1,ai-7,mastra-1,llamaindex-0.12
files: integration/nextjs.test.ts
steps:
- uses: actions/checkout@v7.0.1
- uses: actions/setup-node@v5
with:
node-version: ${{ matrix.node-version }}
cache: npm
cache-dependency-path: |
sdk/typescript/package-lock.json
sdk/typescript/integration/fixtures/*/package-lock.json
- name: Install dependencies
uses: nick-fields/retry@v4
with:
max_attempts: 3
timeout_minutes: 5
command: cd sdk/typescript && npm ci --no-audit --no-fund
# `test:integration` builds, packs, `npm ci`s this shard's fixtures and
# runs its files; the fixture installs are the network-heavy part, so
# retry them as a whole rather than failing the job on a registry blip.
- name: Test against real releases (${{ matrix.shard }})
uses: nick-fields/retry@v4
env:
FAILPROOFAI_IT_FIXTURES: ${{ matrix.fixtures }}
with:
max_attempts: 2
timeout_minutes: 25
command: cd sdk/typescript && npm run test:integration -- ${{ matrix.files }}
test:
runs-on: ubuntu-latest
# The retry above nominally allows 3 attempts x 10 minutes. Capping the job
# at 20 is deliberate: a test job deep into a retry storm is a failure worth
# surfacing, not one worth waiting out.
timeout-minutes: 20
strategy:
fail-fast: false
matrix:
env-config:
- { name: default, env: {} }
- { name: log-debug, env: { FAILPROOFAI_LOG_LEVEL: debug } }
- { name: hook-log-file, env: { FAILPROOFAI_HOOK_LOG_FILE: "1" } }
env:
FAILPROOFAI_TELEMETRY_DISABLED: "1"
NEXT_TELEMETRY_DISABLED: "1"
steps:
- uses: actions/checkout@v7.0.1
- uses: oven-sh/setup-bun@v2
with:
bun-version: latest
- uses: actions/cache/restore@v6
with:
path: ~/.bun/install/cache
key: bun-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-${{ runner.os }}-
- name: Install dependencies
uses: nick-fields/retry@v4
with:
max_attempts: 3
timeout_minutes: 5
command: bun install --frozen-lockfile --ignore-scripts
# Pinned, because vitest runs under Node and this job had been taking
# whatever the runner image happened to ship. That is how the suite came
# to pass here and fail on contributors' machines: from Node 24 on, Node
# supplies its OWN `localStorage` global, which only works with
# `--localstorage-file` and otherwise shadows jsdom's with `undefined` —
# fifteen failures in `project-list.test.tsx`, green in CI. The polyfill
# in `__tests__/setup.ts` is what actually fixes it; this pin is what
# keeps CI from silently drifting onto a different runtime again.
- uses: actions/setup-node@v7
with:
node-version: "22"
# The custom-policy loader tests resolve `import ... from 'failproofai'`
# inside a generated .mjs through findDistIndex(), which needs a real
# dist/index.js on disk — so with `--ignore-scripts` above, the `prepare`
# hook is no longer there to have built it. Three milliseconds of bun
# bundling, against the ~28s full Next build it replaces.
#
# This is also why running the whole suite is a bad way to check the
# dependency: some earlier test writes that file, so the loader tests pass
# on ordering alone when run together and fail when run on their own.
- name: Build the policy-loader fixture
run: bun build src/index.ts --outdir dist --target node --format cjs
- name: Test (${{ matrix.env-config.name }})
uses: nick-fields/retry@v4
env: ${{ matrix.env-config.env }}
with:
max_attempts: 3
timeout_minutes: 10
command: bun run test:run
build:
runs-on: ubuntu-latest
timeout-minutes: 15
env:
FAILPROOFAI_TELEMETRY_DISABLED: "1"
NEXT_TELEMETRY_DISABLED: "1"
steps:
- uses: actions/checkout@v7.0.1
- uses: oven-sh/setup-bun@v2
with:
bun-version: latest
- uses: actions/cache/restore@v6
with:
path: ~/.bun/install/cache
key: bun-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-${{ runner.os }}-
- name: Install dependencies
uses: nick-fields/retry@v4
with:
max_attempts: 3
timeout_minutes: 5
command: bun install --frozen-lockfile --ignore-scripts
- name: Build
uses: nick-fields/retry@v4
with:
max_attempts: 3
timeout_minutes: 10
command: bun run build
docs:
runs-on: ubuntu-latest
timeout-minutes: 10
steps:
- uses: actions/checkout@v7.0.1
with:
# Depth 2 so the gate below can diff the PR's merge commit against
# its first parent. See rust-quality for the full reasoning.
fetch-depth: 2
# This job runs `npm install -g mintlify`, which executes package
# lifecycle scripts. The default would leave GITHUB_TOKEN sitting in
# .git/config where any of them could read it — the same reasoning
# rust-quality and build-daemon already carry, and this is the job
# where it bites hardest.
persist-credentials: false
# Same shape as rust-quality's gate, same reason: this job installs a
# global npm package and a full dependency tree to validate documentation
# that most pull requests do not touch. On `push` it always runs.
- name: Detect docs changes
id: docs
run: |
if [ "${{ github.event_name }}" != "pull_request" ]; then
echo "present=true" >> "$GITHUB_OUTPUT"
exit 0
fi
# bun.lock alongside package.json: `validate:mdx` runs a bun script,
# so a lockfile-only change moves the dependency graph it parses with.
if git diff --name-only HEAD^1 HEAD | grep -qE '^(docs/|scripts/validate-mdx\.ts$|package\.json$|bun\.lock$|\.github/workflows/ci\.yml$)|\.mdx$'; then
echo "present=true" >> "$GITHUB_OUTPUT"
else
echo "present=false" >> "$GITHUB_OUTPUT"
echo "No docs-affecting paths in this PR — skipping validation."
fi
- if: steps.docs.outputs.present == 'true'
uses: oven-sh/setup-bun@v2
with:
bun-version: latest
- if: steps.docs.outputs.present == 'true'
uses: actions/cache/restore@v6
with:
path: ~/.bun/install/cache
key: bun-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-${{ runner.os }}-
- name: Install dependencies
if: steps.docs.outputs.present == 'true'
uses: nick-fields/retry@v4
with:
max_attempts: 3
timeout_minutes: 5
command: bun install --frozen-lockfile --ignore-scripts
- if: steps.docs.outputs.present == 'true'
uses: actions/setup-node@v7
with:
node-version: 22
# Pinned to match translate-docs.yml, which has always pinned it. Floating
# here meant an upstream Mintlify release could turn the PR gate red on a
# branch that changed nothing.
- name: Install Mintlify CLI
if: steps.docs.outputs.present == 'true'
run: npm install -g mintlify@4.2.680
# Validates docs.json structure + nav-link resolution.
- name: Validate docs config
if: steps.docs.outputs.present == 'true'
working-directory: docs
run: mintlify validate
# Parses every MDX page with the same engine Mintlify runs at deploy
# time. `mintlify validate` does NOT parse page content, so a syntax
# error (e.g. a translation that breaks a JSX tag or injects a `{#id}`
# heading anchor) passes that step but fails the post-merge deploy.
# This catches it on the PR instead. It also verifies every image
# reference resolves on disk — that class breaks nothing at build time,
# it just renders as a broken image, so nothing else in CI watches it.
- name: Validate MDX pages parse and image references resolve
if: steps.docs.outputs.present == 'true'
run: bun run validate:mdx
test-e2e:
runs-on: ubuntu-latest
timeout-minutes: 15
env:
FAILPROOFAI_TELEMETRY_DISABLED: "1"
NEXT_TELEMETRY_DISABLED: "1"
steps:
- uses: actions/checkout@v7.0.1
- uses: oven-sh/setup-bun@v2
with:
bun-version: latest
- uses: actions/cache/restore@v6
with:
path: ~/.bun/install/cache
key: bun-${{ runner.os }}-${{ hashFiles('bun.lock') }}
restore-keys: bun-${{ runner.os }}-
- name: Install dependencies
uses: nick-fields/retry@v4
with:
max_attempts: 3
timeout_minutes: 5
command: bun install --frozen-lockfile --ignore-scripts
# The vitest e2e suite spawns bin/failproofai.mjs with
# FAILPROOFAI_DIST_PATH pointed here, so it needs dist/index.js. cli.mjs
# and worker.mjs are for the standalone __tests__/e2e/layout/*.sh
# fixtures, which are not in the vitest `include` and so were quietly
# relying on the `prepare` hook having built them during install. Two more
# bun bundles, ~50ms, and they stay runnable from a CI checkout.
- name: Build E2E fixtures
run: |
bun build src/index.ts --outdir dist --target node --format cjs
bun run build:cli
bun run build:worker
- name: E2E Hook Tests
uses: nick-fields/retry@v4
with:
max_attempts: 2
timeout_minutes: 10
command: bun run test:e2e