Skip to content

Commit f2aa5b4

Browse files
NiveditJainclaudechhhee10
authored
[sdk/typescript] Add @failproofai/sdk, the TypeScript telemetry SDK (#830)
* feat(sdk): add @failproofai/sdk, the TypeScript telemetry SDK The counterpart to the Python SDK at sdk/python: the same 15 events, the same wire format, the same spool directory, the same Evaluator v2 protocol. A fleet running Node agents and Python agents now writes into one pipe, and the dashboard cannot tell which wrote what. Three surfaces, mirroring the Python ones: * Scopes — session(), agent(), toolCall(), with identity carried on AsyncLocalStorage. Each has a callback form and a `using`-compatible .open(). A synchronous body stays synchronous; wrapping every call in a promise would break the one case that genuinely cannot await. * Adapters — instrument() wires LangChain.js/LangGraph.js, the Vercel AI SDK, Mastra and LlamaIndex.TS. * event.* — the 15 methods, with the same validation. The promoted columns are checked at the boundary because ingest answers 200 OK and stores NULL for a value it cannot read, so the alternative is a dashboard that is quietly missing rows. Two places the port deliberately diverges, both because the language differs rather than because the design does: * The Vercel AI SDK exports functions from an ES module, and an ES module namespace is immutable — there is nowhere to patch. It is served by the two extension points that SDK documents: an OpenTelemetry-shaped tracer for experimental_telemetry, and a LanguageModelV2Middleware. Using both records each call once. * Managed evaluator source is PARSED AND INTERPRETED, not eval'd and not handed to node:vm. Python's AST allowlist plus eval does not transfer: x["constructor"]["constructor"]("…")() reaches arbitrary code through a key computed at runtime, which no source-level check can see, and a vm context has its own Function. Every property read goes through one function that checks the actual key at the moment of the read. The worker_threads sandbox around it — V8 heap limits, a wall-clock terminate(), a bounded result, a concurrency cap — is the RESOURCE bound, and it fails closed: no sandbox means refusal. Zero runtime dependencies, enforced by a test and by the build. Dual ESM + CommonJS, Node >= 20.9. 242 tests in the package, plus an 18-case pipeline test in __tests__/ci/. Infrastructure: a CI job across four Node majors that smoke-tests the packed TARBALL (both module systems, --omit=peer, events read back off disk, and the evaluator sandbox resolved through the package's own export — the one part that cannot be exercised from inside this repo); a release workflow with the same preflight/build/publish/verify/bump shape as its Python sibling; the new lockfile added to the OSV scan; and sdk/typescript excluded from the root tsconfig and eslint config, so this project's dependency tree cannot decide whether that package's zero-dependency claim holds. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BJ3C2FkvrjzkxNUqGxvgSW * docs: add the TypeScript SDK reference page and cross-link it The Python custom-agents reference is the page a customer lands on from PyPI, and it was the only place either SDK was documented. A TypeScript reference sits beside it now, registered in the English navigation; the translate pipeline picks up the other fourteen locales on its next run, which is what it is for. The cross-link runs both ways, and both sides say the same thing in the same place: the two SDKs write the same events into the same spool, so a fleet with Node agents and Python agents produces one set of sessions, not two. That is the fact a reader needs before they start choosing, and it does not appear anywhere the choice is actually made otherwise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BJ3C2FkvrjzkxNUqGxvgSW * fix(sdk): stop the TS SDK's tests inheriting the monorepo root's tooling Two failures the isolated CI job found and a local run structurally cannot, because both only appear once this package stands on its own. **Vite read the ROOT's postcss.config.mjs.** It searches upward from the project root for a PostCSS config and finds the one at the repository root, which requires `@tailwindcss/postcss` — a root devDependency that is deliberately absent from this package's node_modules. Vite then dies before a single test runs. Locally the root's node_modules is present and Node's resolution walks up into it, so the search succeeds and nothing looks wrong; in CI, where the isolation is real, it is fatal. An inline empty `css.postcss` turns the search off. This package has no CSS at all, so the only thing that search can do here is find somebody else's tooling. Verified by hiding `@tailwindcss/postcss` from the root node_modules and re-running: 242 passed. **vitest 2.1.9 carried seven fixable advisories.** The Supply Chain gate flagged vite 5.4.21 and esbuild 0.21.5 underneath it, one Critical and one High. osv-scanner.toml says to prefer fixing over ignoring, so this moves to vitest 5 — the major the root project already uses — which pulls vite 8.3.0 and drops the vulnerable esbuild entirely. The Vitest 4 pool rework replaced `poolOptions.forks.singleFork` with the top-level `fileParallelism`, which says the same thing more plainly: run files one at a time, so a test never observes another file's patched prototype or competes with itself for the sandbox semaphore. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BJ3C2FkvrjzkxNUqGxvgSW * ci: prove the TS SDK's Node floor with the artifact, not the test runner The 20.9 leg failed to start: vitest 5 pulls vite 8, which pulls rolldown, which imports `styleText` from `node:util` — added in Node 20.12. Nothing about the package needs it; 20.x, 22.x and 24.x all passed. So the leg says what is actually true. The TEST RUNNER's floor is not the PACKAGE's floor, and the honest way to hold the floor at 20.9 is to prove it with the thing a consumer receives: the floor leg builds, packs, installs the tarball and runs it, and skips the suite the runner cannot start there. The alternative — pinning vitest back to something a 2023-era Node can load — is the tail wagging the dog. It means carrying the CVEs vitest 5 fixed (one Critical, one High, flagged by the Supply Chain gate on the previous push) so that a test runner can start on a release nobody runs the tests on. `__tests__/ci/ts-sdk-pipeline.test.ts` asserts both halves: the floor is in the matrix, at least one leg runs the suite, Build/Pack/smoke are ungated, and exactly Typecheck/Lint/Test carry the gate. A later edit that quietly ungates the suite, or drops the floor leg's smoke test, fails there rather than in a release. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BJ3C2FkvrjzkxNUqGxvgSW * test(sdk/typescript): add a real-framework integration suite; patch the copy the app loads The adapters were only ever tested against src/ with no framework installed, so nothing proved an adapter reached a real framework. They didn't: every adapter resolved its framework with createRequire, which names the CommonJS build of a dual-published package, and patched that copy. An ES-module app (the default for new TypeScript projects) loads the other copy, so instrument() reported success and recorded nothing. integration/ installs real framework releases from per-fixture lockfiles, extracts the PACKED tarball into each, and runs one agent.ts as both ESM and CJS. Expected traces are the Python SDK's golden output for the same program. compat.requireModuleCopies() now returns the copy the application's imports reach (by the entry point's module system, resolved through the package's exports map with ESM conditions) plus the CommonJS copy if something already required it, and never loads a copy speculatively. The LangChain adapter now patches every copy; its mapping is still the pre-parity one. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ci(sdk/typescript): run the framework adapters against real releases; test copy selection A `failproofai-ts-sdk-integrations` job, the counterpart of the Python SDK's `failproofai-sdk-integrations`: packs the SDK, installs every fixture from its lockfile, and runs each agent as ESM and CJS on Node 20 and 24. test/copies.test.ts pins which copy of a dual-published package requireModuleCopies() returns, from real ESM and CJS entry points against a fake dual package on disk: the ESM build for an ES-module app, the CJS build for a CommonJS app, both when the CJS copy was already required, and never a speculatively loaded second copy. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): ship CommonJS declarations and node10 subpath types The exports map handed the ESM .d.ts files to `require`, so TypeScript read every CommonJS import of the package as an ES module: TS1479 on every import under `module: node16` (and `nodenext` before TS 5.8), and attw "Masquerading as ESM" for node16-from-CJS on every entry plus node16-from-ESM on ./sandbox-worker. `moduleResolution: node` (node10, what `module: commonjs` implies) resolved the root and no subpath at all. - the CJS build now emits declarations; under dist/cjs's `"type": "commonjs"` they are CommonJS declarations - exports use nested import/require conditions, each with its own `types`; ./sandbox-worker (CommonJS only) points its types at dist/cjs - `types` points at the CJS declarations and `typesVersions` maps each subpath for node10; exports-aware resolvers ignore it - finalize-build walks nested conditions, refuses a half whose types and JavaScript live in different directories, and checks typesVersions integration/types.test.ts typechecks a consumer of every entry point under six consumer tsconfigs on TS 5.9.3 and 5.4.5 (the minimum: the oldest with `module: preserve`), skipLibCheck off, and runs attw over the packed tarball. test/packaging.test.ts learns the nested condition shape and asserts it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): make the Vercel AI SDK adapter record real ai 4-7 traces Verified against real `ai` releases, the adapter was broken in six ways: instrument("ai") always threw (setGlobalTracerProvider called unbound) and its "don't overwrite a provider" check read the always-present proxy; startActiveSpan auto-ended spans, so streamText lost the half the SDK ends itself (endWhenDone:false); the README call sites did not typecheck on ai 5/6; ai 7 was unsupported, silent, and ERESOLVE'd on install; v3 usage objects and {unified,raw} finish reasons lost tokens and wrote objects into stop_reason; and a bare wrapModel call was dropped or pinned to a phantom "main" agent. The adapter now follows the Python mapping rule: one operation = one agent named by functionId (never an id), each model step a request/response pair with integer tokens and a string stop reason, each tool a pair on the model's toolCallId, a failure recorded once where it happened. v4-v6 attach through a structural OTel tracer that never ends spans for the caller; v7 through its Telemetry integration interface (per call via telemetry(), or globalThis.AI_SDK_TELEMETRY_INTEGRATIONS via instrument()). telemetry() returns both, and v6's callId-less integration events are ignored, so one call site works on every major and nothing records twice. A bare wrapModel call becomes its own run named after the model unless an agent() owns it. Public types are structural and typecheck against real ai 4/5/6/7 types. Peer range widened to ai >=4.0.0 <8. integration/ adds ai-4..ai-7 fixtures (pinned lockfiles, ESM + CJS, a nodenext typecheck of the README call sites). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): make the LangChain adapter draw the Python adapter's trace The LangChain.js / LangGraph.js adapter recorded every LangGraph node as a nested agent, every tool call under its run id, a run-level `error` event per layer a failure unwound through, and nothing at all for sessions keyed by metadata. It is now a port of sdk/python's integrations/langchain.py, held to that adapter's golden output by the real-framework suite: - root run = agent; a LangGraph node = hook_triggered/hook_completed with trigger_event "graph_node" (Python's `_node_of` rules); a compiled subgraph = nested agent "root/node"; machinery emits nothing unless `includeChains` - tool_use/tool_result carry the MODEL's tool_call_id (core 1.x passes it; on core 0.3 it is recovered from the ancestor's tool_calls) - failure is carried by model_response.error, hook_completed failed and agent_end failed; a standalone `error` only when no span owned it - interrupt() -> human_wait + agent_pause, Command resume -> agent_resume + human_input on the same agent, including a resume taken by ANOTHER process (the pause id rebuilt from the checkpoint namespace, as LangGraph derives it) - session order: sessionId option, metadata.failproofai_sdk_session_id, ambient scope, session_id/conversation_id/thread_id, root run id - streaming folded into model_response (fw_streamed, fw_chunks, fw_ttft_ms); .batch() roots stay separate; aborts end `cancelled`, including the root LangGraph.js 1.x abandons without an end callback - langchainHandler() works without instrument(), never double-records with it, and uninstrument() closes open runs as cancelled - the handler is awaited (awaitHandlers) so events keep the caller's async context, and configure() receives it as an input handler so the call's metadata is not dropped Adds a langchain-0.3 fixture (@langchain/core 0.3.80, @langchain/langgraph 0.4.10) run over the same expectations, so the declared >=0.3.0 <2 range is tested at both ends. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): make the Mastra adapter record real Mastra runs Verified against @mastra/core 1.68.0 and 0.24.9, the adapter recorded nothing in an ES-module app, never saw a tool, labelled every agent "mastra-agent", collapsed a multi-step loop into one model pair with no model or tokens, ended streams before they ran, and dropped a bare wrapTool() call. Rewritten on patch points every run goes through, on every copy of @mastra/core the app loads (compat.requireModuleCopies): - Agent.generate/stream (+ legacy/VNext): the agent span, named after the agent, nested under an enclosing scope or the delegating agent. stream() ends on Mastra's onFinish/onError/onAbort, not on return. - Agent.resolveModelConfig: the resolved model goes back behind a proxy observing doGenerate/doStream, so each LLM step is its own model_request/model_response with model id, tokens, stop reason. - Agent.convertTools: wraps the loop's converted tools, so every tool (including ones built before instrument()) records with the model's toolCallId, attributed to the agent whose loop called it. - Run._start/_resume and DefaultExecutionEngine.executeStep: workflow runs as agents, steps as workflow_step hooks; Mastra's own internal workflows (the agentic loop) are skipped. Mastra 1.x's ObservabilityExporter was considered and rejected: it only fires for agents registered on a Mastra instance with @mastra/observability configured, so a bare Agent records nothing. Peer floor raised 0.10.0 -> 0.20.0, where resolveModelConfig and Run._start first exist; earlier lines route model calls through the AI SDK and would record no steps. Adds integration/mastra.test.ts with mastra-1 (1.68.0) and mastra-0 (0.24.9) fixtures, ESM and CJS, and unit tests for the readers. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): make the LlamaIndex.TS adapter record real runs The adapter recorded nothing in an ES-module app, split every legacy LLMAgent run into two sessions with an agent left open, dropped streamed token usage and every model name, and dropped bare llm.chat() calls. Nothing recorded an agent().run() at all: @llamaindex/workflow emits no run or step events on the callback bus. It now attaches in two places. The callback bus (@llamaindex/core/global, subscribed on each copy the app uses, via compat.requireModuleCopies) carries model, tool, retrieval, query and legacy-runner events. AgentWorkflow's runStream is patched to open the run and attach to the context's __internal__call_context / __internal__call_send_event middleware hooks, the surface workflow-core's own withTraceEvents/withState middleware uses, so steps become hooks and the stop event ends the run. Bus events are placed by an AsyncLocalStorage frame bound around each step, or by LlamaIndex's own EventCaller chain (legacy runners, query engines, providers' chat). Same tree as the Python adapter: run = agent named after the agent, multi-agent handoff = nested agents, step = workflow_step hook, bare call = its own run. Streamed usage is read off the chunk that carries it; the model name comes from the caller chain or the agent the step names. The floor moves from 0.9.0 to 0.11.4, the first llamaindex on the workflow 1.1 runtime; 0.9-0.11.3 ship workflow 1.0, whose AgentWorkflow has no runStream. Integration fixtures pin 0.11.4 (the floor) and 0.12.1. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(sdk/typescript): state the tested framework ranges and the Python mapping The adapter table claimed LangGraph nodes become agent spans, Mastra's createTool is patched and LlamaIndex is never patched — none true any more — and stated no supported versions at all. It now names the range each adapter is tested against, the mapping it shares with the Python SDK, how the ESM/CJS dual-package case is handled and where it cannot be (bundled output), the ai 7 `telemetry:` spelling, and langchainHandler() without instrument(). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): bound what a long-running process retains; find the app from its entry From an adversarial review, each reproduced first: - RunTracker.links was never pruned when a run ended, so it filled to its 10k FIFO cap with finished runs and then evicted the links of LIVE ones: a LangChain model call that outlived ~10k other runs lost its model_response to "could not resolve a session". Links now go when their run closes (emit() drops them after a closing event, endAgent after agent_end, LangChain after every end path), and a re-link refreshes an entry's age. - A LangGraph run paused on a human and resumed by another worker was held open here forever, taking a tracker slot live runs need, and every callback scanned all open agents. isOpen() is O(1); paused sessions are forgotten (never closed — the other worker closes them) after PAUSED_SESSION_TTL_MS or when they fall off the session cap. A late resume takes the cross-worker path. - Framework resolution was anchored at process.cwd() alone: a service started from / found no framework, and in a monorepo it patched the hoisted root copy instead of the app's own. It now resolves from the entry script's directory and the working directory, preferring the entry unless it is a launcher inside node_modules, and a copy already required wins. - `@internal` test hooks no longer ship in the published declarations (stripInternal). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): stop the Mastra adapter recording after uninstrument() Model proxies and tool wrappers built while instrumented called ensureTracker(), which lazily built a FRESH tracker after teardown. A stream in flight at uninstrument() therefore kept emitting model/tool events - attributed to a phantom "main" agent inside any open scope, plus a second agent_end - and anything an instrumented run had built stayed live forever. - One tracker per installation, created by install() and dropped by uninstall(); an `enabled` flag checked on every recording path. Every proxy, wrapper, span and frame remembers the tracker it began on and becomes a pass-through once that tracker is gone, including after a later instrument(). - uninstall() switches off first, then closes what is open: model steps (stop_reason "cancelled"), tool calls and workflow steps (all marked fw_incomplete), then agents "cancelled". - Open model/tool/step spans are tracked in bounded registries that every end path prunes (success, error, abort, cancel); a never-consumed stream is held only until teardown closes it. - A model proxy or tool wrapper from an earlier run is re-wrapped for the current run instead of being skipped as already wrapped. - instrument()'s captureLimit is honoured even when wrapTool() recorded first (it used to be ignored). wrapTool() keeps working without instrument() on its own self-contained tracker. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): Mastra calls cut off by uninstrument() stop as "incomplete", like LangChain's Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): keep concurrent LlamaIndex runs on one shared object apart Runs were keyed by the object that owns them (the first computedCaller), and a query engine or LLMAgent built once and shared by every request is the same object for every call: request B's query-start found A's run, treated itself as nested and recorded nothing; B's retrieval landed in A's session and B's answer was never recorded. Legacy chats nested B under A and sent A's later model calls to B. Query and legacy runs are now keyed by LlamaIndex's own EventCaller - a fresh object per @wrapEventCaller invocation, bound in LlamaIndex's AsyncLocalStorage and chained through .parent - and an event belongs to a run only when that exact invocation is on its chain. This is used instead of prototype-patching query/chat because @wrapEventCaller binds the method onto each INSTANCE at construction, so a prototype patch would miss every engine built before instrument(). A bus with no EventCaller falls back to owner matching, where a run owned by the object starting a new run is never taken as its parent. Also: - attach() throws while an install is live instead of orphaning its bus subscriptions and runStream patch; documented @internal. - Every tracker link a run makes (run key, leaf keys, workflow step keys) is unlinked on every end path, so completed runs leave no residue and the tracker's FIFO cap never evicts a live run's link. Adds RunTracker.unlink(). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): stop instrument("ai") taking the global OTel slot; close cancelled/errored streams instrument("ai") on ai 4-6 registered FailproofTracer as the process-wide OpenTelemetry tracer provider whenever the slot was empty. OpenTelemetry refuses every later registration, so a customer's NodeSDK.start() further into startup was silently refused and their http/pg/Next.js spans went to a tracer that exports nothing. Taking the slot is now opt-in (instrument("ai", { registerGlobalTracer: true })); by default instrument("ai") on 4-6 logs one warning naming the call-site telemetry()/wrapModel() paths, and registerGlobalTracer: false silences it. The v7 global integration list is additive and per-call integrations replace it, so it is kept. middleware()/wrapModel() observed streams with pipeThrough(TransformStream), whose flush never runs on a consumer cancel or a stream error: the model call stayed open forever and a standalone ai-model run never got agent_end. It now uses a pull-based re-stream (core.observeStream, extracted from mastra.ts, which now calls it): cancel closes the pair with stop_reason "cancelled" and ends the run cancelled (the cancel reaches the provider stream); an error closes it with the error and ends the run failed. Every tracer span, v7 call/model/tool and middleware request now unlinks its RunTracker parent link when its last event is out, including tools a v7 operation abandoned. Adds RunTracker.unlink(). 20k completed operations leave no residue and no longer evict a live run's links. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(sdk/typescript): changelog for the long-running-server fixes and the instrument("ai") opt-in Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): record LangChain roots started through a nested @langchain/core; cover every common LangChain.js surface Coverage sweep of the LangChain.js adapter against real releases, both module systems, both ends of the peer range: - langchain-0.3 + langchain-1: LCEL (prompt|model|parser), RunnableSequence, RunnableParallel, RunnableLambda, a custom retriever, a VectorStore .asRetriever(), a RAG chain, .batch() x3, .streamEvents() v2, chain .stream(), withStructuredOutput(zod), bindTools with tool_choice — each under instrument() AND through an explicit langchainHandler(). Expected traces are the Python adapter's for the equivalent program. - langchain-1: the v1 `langchain` package's createAgent (1.5.12), plain, streamed, explicit-handler and with middleware (agent-v1.ts). - langchain-dup-core (new fixture): an app on core 1.2.12 whose vendored provider pins core 0.3.80, so npm nests a second copy under it. Bugs found and fixed: - Runs a nested @langchain/core copy started as ROOTS (a provider's model, tool, retriever or runnable invoked directly) were recorded nowhere under instrument(): only the app's copy of CallbackManager was patched. LangChain has no cross-copy hook (registerConfigureHook keys on a module-private Symbol), so install now finds nested copies on disk (node-require.nestedCopies, incl. the pnpm store) and patches each in-range copy's reachable build(s). - createAgent's model node returns a LangGraph Command; the payload view handed it to truncate() whole, so its messages were captured as LangChain's {lc, type: "constructor", id, kwargs} envelope. Harness: additive — transpile every agent*.ts program of a fixture, and runAgent takes an optional program name. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): AI SDK coverage sweep — embeddings, abandoned streams, tool-call content Walk every commonly used ai 4–7 surface against the packed SDK (new surfaces.ts per fixture, ESM + CJS): agent classes, embed/embedMany, structured output, tool features, stream consumption styles, reasoning models and concurrency. Fixes what that exposed: - embed/embedMany inside an agent() scope or a tool are model calls of the enclosing agent, not nested ai.embed agents (bare calls stay their own run) - v4–v6: an aborted stream closes its open model/tool spans as cancelled and ends the agent cancelled (was: success with an unpaired model_request) - v4–v6: a stream the SDK never ends (client disconnect, never read, v4 mid-stream provider error) is closed when its root span is collected (FinalizationRegistry, fw_abandoned) instead of staying open forever - v4–v6: tool calls in model_response content use one shape with parsed input (was v4 {toolCallType,args:"<json>"} / v5-6 input as a JSON string) - v7: an aborted model call closes as cancelled (was "incomplete") and a failed/aborted one keeps its model id Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): Mastra coverage sweep — memory threads, networks, suspend/resume, tripwires Integration cases (ESM + CJS, @mastra/core 0.24.9 and 1.68.0) for every commonly used Mastra surface: agents/workflows fetched from a Mastra instance, 10 concurrent runs of one Agent, @mastra/memory threads, agent networks, agent-in-tool delegation, branch/parallel/dowhile/foreach/nested workflows, createStep(agent), run.stream(), suspend/resume, input/output processors and tripwires, MCP tools over a local stdio server, structured output (incl. a second structuring model), maxSteps with a failing tool, and streamed usage from an OpenAI-compatible endpoint. Bugs they exposed, fixed in the adapter: - a memory thread was not the session: every turn of one conversation was a new session, and 1.x's memory: { thread } never reached fw_thread_id; - agent.network() shattered into one root session per routing decision ("Routing Agent" x2 + the delegate), and on 0.x recorded Mastra's own network workflows plus the delegate's internal loop steps as hooks; - workflow suspend/resume was two unrelated sessions with no HITL events; now human_wait + agent_pause ... agent_resume + human_input on one span, with deterministic ids and a cross-process resume that closes them; - a stream() blocked by a processor tripwire left its agent open forever. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): LlamaIndex coverage sweep — chat engines, retriever names, plain workflows Real-framework cases (llamaindex 0.11.4 + 0.12.1, ESM + CJS) for every commonly used LlamaIndex.TS surface, checked against the Python adapter's golden traces. Three bugs they exposed: - A chat engine (Simple/Context/CondenseQuestion) dispatches nothing of its own, so its retrieval and model call were separate root runs in separate sessions, named after the model. Its run boundary is now read from LlamaIndex's own EventCaller storage (an own `run` on that one object): a top-level invocation is one run named after its class, ended when it returns. The same signal closes a model call, query or legacy task that THREW, which previously stayed open until the reaper. - Retrievals were named "retriever"; Python names them after the retriever class. BaseRetriever.prototype.retrieve now binds the retriever for its retrieve-start. - createWorkflow() workflows recorded no run and no steps. On workflow-core >=1.1 its exported AsyncContext.Variable sees every step handler; a context's burst of steps is now a "Workflow" run with hook pairs. Also covered: multiAgent() with 3 agents, responseFormat, FunctionTool.from, QueryEngineTool, parallel tool calls, a mid-stream provider error (run closes failed despite LlamaIndex's unhandled rejection), and 10-way concurrency on a shared agent, query engine and chat engine. Unit tests for each fix and for @llamaindex/openai's include_usage stream chunk. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): flush() waits for events emitted during an in-flight flush flushNow() — what the public flush() awaits — handed back a flush that was already running, which had drained the queue before any event emitted since. So `await failproofai.flush()` could resolve with the newest events still in memory: exactly the emit-flush-exit case it exists for. It showed up as a ~1-in-6 flaky unit test. It now flushes twice: the second pass drains what the first could not see, or joins a newer flush that already did. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(sdk/typescript): run every framework fixture under Bun and Deno; fix Deno ESM apps recording nothing Deno sets require.main to null (not undefined) for an ES-module entry, so entryIsCommonJs() read every ES-module Deno app as CommonJS and instrument() patched the frameworks' CommonJS copies while the app ran the ES-module ones: LangChain, Mastra and LlamaIndex reported success and recorded nothing. Adds a pinned runtimes fixture (bun 1.4.2, deno 2.9.6 from npm, located by the harness), bun-*/deno-* formats, a Node-parity suite over all 184 framework scenarios per runtime, a core smoke of every scope and event method, the short-lived-handler (Lambda) cases, and Deno npm: specifiers. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): instrument() records under Next.js serverExternalPackages; importing the SDK no longer breaks an Edge build Next.js: next start is a CommonJS launcher, but Next loads every server-external package with import() (Turbopack and webpack alike), so the app runs the frameworks' ES-module copies. The SDK read the launcher as a CommonJS app and patched the CommonJS copies: instrument() in instrumentation.ts reported success for LangChain, Mastra and LlamaIndex and recorded nothing. Inside a Next server (NEXT_RUNTIME set) the ES-module copy is now the one patched (appImportsReachCommonJs in node-require.ts, used by requireModuleCopies). Edge: an Edge-runtime route importing the SDK failed next build outright ("Native module not found: node:fs"). The package now maps the edge-light, workerd, worker and browser conditions to a no-op build (src/edge/) with the same public surface that logs one line on first use; types stay the real ones. Adds a Next.js 16 App Router fixture (one route per framework, a server action, instrumentation.ts, an Edge route) built four ways - Turbopack/webpack x default(bundled)/serverExternalPackages - with each route's trace compared to the same program's trace under plain Node. Drops the short-lived-handler cases from the runtimes suite (not a deployment target). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(sdk/typescript): reconcile the parallel coverage sweeps Each sweep was green alone; merged, three things disagreed: - the edge no-op build's parity test did not know two helpers the ai and llamaindex sweeps export for their unit tests (toolCallsOf, eventCallerStorage) — both marked @internal and listed as such; - runtime parity compared event ORDER within a session for the Mastra sweep's `concurrent` scenario, whose ten runs share one session and interleave by scheduling on every runtime — such scenarios now compare the multiset; Mastra 0.x `tripwire-output`, a documented limit its own suite skips, is skipped here too; - Deno's node:http keeps idle keep-alive sockets open through close(), so the loopback provider in `usage-openai-compatible` held the process past its timeout — the fixture now closes them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(sdk/typescript): withFailproofai() for Next.js, and a warning when Next bundles a framework `next build` bundles LangChain, Mastra and LlamaIndex into its server output by default, so `instrument()` patched the node_modules copy the app never runs, reported success, and recorded nothing. `@failproofai/sdk/next` exports `withFailproofai(nextConfig)`, which adds the packages the adapters need (and the SDK itself, so instrumentation.ts and the routes share one copy) to `serverExternalPackages`, keeping the app's own list, any config form (object, sync or async function), and leaving out anything the app lists in `transpilePackages`. `instrument()` on a Next.js Node server now warns once per adapter that cannot reach a bundled copy — LangChain, Mastra, LlamaIndex, never the AI SDK — unless the app is configured: the list the wrapper records when Next evaluates the config, a standalone server's resolved config, or a hand-set FAILPROOFAI_NEXT_EXTERNALS=1. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(sdk/typescript): build each Next.js variant from scratch `npm pack` stamps every file with one fixed mtime, so when the harness put a newly packed SDK into the fixture, webpack's persistent cache (kept in the distDir) saw it as unchanged and bundled the PREVIOUS run's SDK. The webpack variants were silently testing stale SDK code; Turbopack hashes content and was unaffected. Found because the webpack build printed none of the new Next.js warnings while Turbopack printed all three. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * ci(sdk/typescript): shard the real-framework suite; guard the shard lists Frameworks at both ends of their ranges, Bun and Deno parity over every fixture, and five real `next build`s are ~23 CPU-minutes — too much for one runner inside a timeout. The job is now a matrix: frameworks (Node 20 and 24), runtimes, and nextjs, each installing only its own fixtures and running only its own files. The shard lists are hand-maintained, so a new integration file left out of every shard would never run in CI. __tests__/ci/ts-sdk-integration-shards.test.ts fails when a file is uncovered or a shard names a file or fixture that does not exist. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * docs(sdk/typescript): Next.js setup, runtimes, per-framework notes and streamed-usage caveats Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * test(sdk/typescript): cover the ./next entry point in the declarations suite The attw check pins the public entry-point list, so the new `./next` export failed it. It is now listed, and the consumer file every resolution mode typechecks imports withFailproofai and proves its result is typed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * feat(sdk/typescript): instrument your own agent — guide and a CI-proven example For a client whose agents are hand-built TypeScript with no framework. The Python SDK's answer carries over one to one: every such agent already has three places (where a run starts and ends, the one function that calls the model, the one function that runs tools), and the scopes at those three are the whole integration, producing the same trace the adapters do. examples/research-agent.ts is the TypeScript twin of the Python SDK's research_agent.py — a real OpenAI tool loop instrumented by hand, pairing model calls on requestId, timing them, closing them on failure, and reusing the model's tool-call ids. The `vanilla` fixture IS that file (a test holds them byte-identical) and runs it as ESM and CJS on the real openai client against a local OpenAI-compatible server: parallel tools, a failing tool, a failing model. README and docs gain a "your own agent — no framework" section. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> * fix(sdk/typescript): run CommonJS evaluator files; name the real ai import in warnings failproofai-evaluator is the ESM build, and a CommonJS evals file builds its Evaluator from dist/cjs, a second copy of the class, so the loader's instanceof check refused it. It now reads a Symbol.for brand both builds stamp. A packaging test runs a CommonJS and an ESM file through the built bin. The ai adapter's warnings told readers to call failproofai.ai.telemetry(), which the root package does not export; they now name the @failproofai/sdk/ai import. * docs(skill): failproofai-sdk covers TypeScript agents and the evaluator worker New references/typescript.md (install, names, adapters, Next.js and bundlers, the three-edit-site wiring for a hand-built loop, shutdown, verification) and references/evaluator.md (writing, running and deploying the eval pod in Python and TypeScript). The description no longer routes to the retired agenteye-evaluator, and agents/openai.yaml takes the policy: nesting the skills repo already carries so the next sync does not revert it. test/skill-snippets.test.ts parses every TypeScript block in the skill and fails when one calls an SDK name that no longer exists. * fix(sdk/typescript): what six user-style runs against Cloud found - Timestamps: events in one millisecond tied, and the dashboard showed a tool_result before its tool_use. The digits below the millisecond are now a per-process sequence (clock.ts), re-anchoring if the wall clock steps back. - Exit: SIGTERM left runs rendered as running forever. The exit hook closes open tools (ProcessExit error) and agents (error + agent_end failed), including adapter-opened runs; a run paused on a human stays open for the process that resumes it. - configure() reached only its own copy of the SDK, so a Next.js route bundled without withFailproofai reported environment dev. environment and baseDir are process-wide now (shared.ts). - An agent() wrapper around a framework agent of the same name joins it instead of opening a self-parented duplicate. - LlamaIndex 0.12 tool errors no longer read 'Error: Error(Error): ...'; error_type is the class for errors whose name is just 'Error'; an unawaited instrument() is warned about by the LangChain adapter. - The vanilla example records the model's tool calls, keeps its error path to the provider call, and survives malformed tool arguments. The skill's TypeScript page covers naming, owning the session id, adapter options, streamed-usage flags, verifying under a running daemon, Next.js configure() placement and the new shutdown behaviour. * fix(sdk/typescript): close every open leaf at exit; a strict-typed vanilla recipe Round two of the user-style tests (six agents, real models, Cloud): - Shutdown closed an adapter's agent but left its open tool, node and model call — and a hand-written event.modelRequest — unclosed. Tool calls, hooks and model calls are now tracked in the event namespace, which every one of them is emitted through, and closed at exit in a 'leaves' phase before any agent; a session paused on a human is skipped. Adapter agents closed at exit also get their ProcessExit error event now. - The unawaited-instrument() warning never fired live: the graph's nodes arrive as roots, not under an unknown parent. A graph node arriving as a root now warns. - The skill's no-framework snippet failed tsc --strict on openai's union tool types, dropped a malformed-arguments call from the trace, and lost the ids linking tool results to their calls. Fixed in the example and the skill, recorded as fw_tool_calls like the adapters, and the skill's block is now type-checked against the real openai client in the vanilla suite. - Docs: several AI SDK calls in one session() are several agents; the Next.js warning prints at boot; ToolLoopAgent naming; LlamaIndex's default temperature; a failed tool does not fail an error-count evaluation. * fix(sdk/typescript): close at exit most-recent-first; name the crash that exited Round three of the user-style tests (killed runs on LangGraph, Next.js, Mastra, LlamaIndex, the AI SDK and a hand-built loop, checked in Cloud): every tool, hook, model call and agent now closed on every exit path. What was still wrong: - Order: tools, hooks and model calls closed before any agent, so a planner's delegate tool ended while the writer sub-agent it started was still running. Owners now hand the exit path what they hold open, and it closes everything in one most-recently-opened-first order. - A crash left 'the process exited (code 1)' and no trace of the exception. The closing messages name it (read via uncaughtExceptionMonitor, which observes without changing how the process dies). - A model call closed at exit carries its duration; a ProcessExit error no longer carries the SDK's own stack as its traceback. - Recipe and skill: tool arguments must parse to an object; the recorded history keeps the name of each tool the model asked for. The skill covers every exit path, the app's own records on SIGTERM, npx swallowing SIGTERM, a missed await passing silently, Next.js request ids and cancellation, and that a killed run's session status is still 'done'. --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> Co-authored-by: chhhee10 <chetanraghuvanshi85@gmail.com>
1 parent aefc062 commit f2aa5b4

209 files changed

Lines changed: 68780 additions & 36 deletions

File tree

Some content is hidden

Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.

‎.github/workflows/ci.yml‎

Lines changed: 251 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -484,6 +484,257 @@ jobs:
484484
AGENTEYE_TESTS_REQUIRE_FRAMEWORKS: "1"
485485
run: uv run pytest tests/integrations -q
486486

487+
# The TypeScript telemetry SDK (`@failproofai/sdk`). Separate from `quality`
488+
# and `test` because it is its own npm package with its own lockfile, its own
489+
# tsconfig and its own vitest config — running it inside the root project's
490+
# jobs would mean the root's dependency tree decided whether this package's
491+
# zero-dependency claim holds.
492+
#
493+
# The matrix is Node's supported majors, not a single version. This package's
494+
# floor is 20.9 and its `using` support, `AsyncLocalStorage.enterWith`,
495+
# `worker_threads` resource limits and `Symbol.dispose` shim all behave
496+
# differently across that range — which is precisely the range a customer's
497+
# agent runs on.
498+
failproofai-ts-sdk:
499+
runs-on: ubuntu-latest
500+
timeout-minutes: 15
501+
defaults:
502+
run:
503+
working-directory: sdk/typescript
504+
strategy:
505+
fail-fast: false
506+
matrix:
507+
include:
508+
# The `engines` FLOOR. This leg proves the published artifact runs on
509+
# it, and deliberately does not run the suite: vitest 5 pulls vite 8,
510+
# which pulls rolldown, which imports `styleText` from `node:util` —
511+
# added in Node 20.12. The TEST RUNNER's floor is not the PACKAGE's
512+
# floor, and the honest way to say so is to keep the floor at 20.9 and
513+
# prove it with the thing a consumer actually gets.
514+
#
515+
# Pinning the runner back to something 20.9 can load is the tail
516+
# wagging the dog: it means carrying the CVEs vitest 5 fixed (one
517+
# Critical, one High) so that a test runner can start on a Node
518+
# release nobody runs the tests on.
519+
- node-version: "20.9"
520+
suite: false
521+
- node-version: "20.x"
522+
suite: true
523+
- node-version: "22.x"
524+
suite: true
525+
- node-version: "24.x"
526+
suite: true
527+
steps:
528+
- uses: actions/checkout@v7.0.1
529+
530+
- uses: actions/setup-node@v5
531+
with:
532+
node-version: ${{ matrix.node-version }}
533+
cache: npm
534+
cache-dependency-path: sdk/typescript/package-lock.json
535+
536+
- name: Install dependencies
537+
uses: nick-fields/retry@v4
538+
with:
539+
max_attempts: 3
540+
timeout_minutes: 5
541+
command: cd sdk/typescript && npm ci --no-audit --no-fund
542+
543+
- name: Typecheck
544+
if: matrix.suite
545+
run: npm run typecheck
546+
547+
- name: Lint
548+
if: matrix.suite
549+
run: npm run lint
550+
551+
# Runs on EVERY leg, including the floor: `tsc` is the thing that produces
552+
# what ships, so "it builds on 20.9" is a claim worth checking there.
553+
- name: Build
554+
run: npm run build
555+
556+
# The sandbox suite needs `dist/` present — the evaluator sandbox is a
557+
# real `worker_threads` entry and cannot load a `.ts` file — and it
558+
# asserts rather than skips when it is missing, so the build above is a
559+
# prerequisite rather than a duplicate.
560+
- name: Test
561+
if: matrix.suite
562+
run: npx vitest run
563+
564+
# Everything above ran against the source tree. These two steps run
565+
# against the ARTIFACT, because the failures they catch — a missing export
566+
# condition, a CommonJS build Node reads as ESM, a `dist/` path the files
567+
# list does not ship — are invisible from inside the package and total
568+
# from outside it.
569+
- name: Pack
570+
run: npm pack --pack-destination /tmp
571+
572+
- name: Smoke-test the packed tarball with no dependencies
573+
run: |
574+
mkdir -p /tmp/ts-sdk-smoke && cd /tmp/ts-sdk-smoke
575+
npm init -y >/dev/null
576+
# `--omit=optional --omit=peer` is the assertion, not an optimisation.
577+
# "Zero dependencies" is the reason this package is safe to drop into
578+
# someone else's agent, so prove it against the built artifact: install
579+
# it with nothing else present, emit real events, and read them back
580+
# off disk.
581+
npm install --no-audit --no-fund --omit=optional --omit=peer /tmp/failproofai-sdk-*.tgz
582+
test ! -d node_modules/@failproofai/sdk/node_modules \
583+
|| { echo "the published package brought transitive dependencies"; exit 1; }
584+
585+
cat > esm.mjs <<'EOF'
586+
import * as fp from "@failproofai/sdk";
587+
fp.configure({ baseDir: process.env.SPOOL });
588+
await fp.agent("smoke", { goal: "ci" }, async () => {
589+
await fp.toolCall("t", { input: { q: 1 } }, () => "ok");
590+
});
591+
await fp.flush();
592+
console.log(fp.version);
593+
EOF
594+
SPOOL=/tmp/ts-sdk-spool node esm.mjs
595+
596+
cat > cjs.cjs <<'EOF'
597+
const fp = require("@failproofai/sdk");
598+
fp.configure({ baseDir: process.env.SPOOL });
599+
fp.event.agentStart({ sessionId: "cjs", goal: "ci" });
600+
fp.flushSync();
601+
console.log(fp.version);
602+
EOF
603+
SPOOL=/tmp/ts-sdk-spool node cjs.cjs
604+
605+
# The evaluator loads from its own subpath, and its CLI has to be
606+
# executable — a `bin` that is not fails on every platform where the
607+
# installer links rather than copies.
608+
node -e "const e = require('@failproofai/sdk/evaluator'); if (typeof e.Evaluator !== 'function') throw new Error('evaluator subpath is broken')"
609+
npx --no-install failproofai-evaluator --help > /dev/null
610+
611+
node - <<'EOF'
612+
const { readdirSync, readFileSync } = require("node:fs");
613+
const { join } = require("node:path");
614+
const dir = "/tmp/ts-sdk-spool/events";
615+
const events = readdirSync(dir)
616+
.filter((f) => f.endsWith(".jsonl"))
617+
.flatMap((f) => readFileSync(join(dir, f), "utf8").split("\n").filter(Boolean))
618+
.map((line) => JSON.parse(line));
619+
const types = new Set(events.map((e) => e.type));
620+
const want = ["agent_start", "agent_end", "tool_use", "tool_result"];
621+
for (const type of want) {
622+
if (!types.has(type)) throw new Error(`the installed artifact never wrote ${type}`);
623+
}
624+
// Emitted by the artifact, so this also proves the wire format
625+
// survived packaging rather than only surviving an in-tree import.
626+
if (!events.every((e) => "environment" in e && "session_id" in e)) {
627+
throw new Error("an event reached disk missing a required field");
628+
}
629+
console.log(`${events.length} events written by the installed artifact`);
630+
EOF
631+
632+
# The evaluator sandbox in an INSTALLED package resolves its worker
633+
# through the package's own `./sandbox-worker` export, with no env
634+
# override in sight. That resolution is the one part of the sandbox that
635+
# cannot be exercised from inside this repository, and a failure in it
636+
# means managed evaluations refuse to run for every customer.
637+
- name: Verify the evaluator sandbox resolves from an installed package
638+
run: |
639+
cd /tmp/ts-sdk-smoke
640+
node - <<'EOF'
641+
const { compileEvaluator, sessionTranscriptFromWire } = require("@failproofai/sdk/evaluator");
642+
const session = sessionTranscriptFromWire({
643+
schema_version: "2",
644+
assignment_id: "a", session_id: "s", session_revision_id: "r",
645+
agent_id: "main", environment: "dev",
646+
started_at: "2026-01-01T00:00:00.000000Z",
647+
ended_at: "2026-01-01T00:01:00.000000Z",
648+
event_count: 1,
649+
events: [{ id: "1", ts: "2026-01-01T00:00:01.000000Z", event_type: "tool_use", payload: {} }],
650+
});
651+
compileEvaluator("EvalResult({ score: Score(session.count('tool_use') > 0 ? 1 : 0) })", { evalKey: "k" })(session)
652+
.then((result) => {
653+
if (result.score.value !== 1) throw new Error(`unexpected score ${result.score.value}`);
654+
console.log("sandbox resolved and evaluated from the installed package");
655+
})
656+
.catch((error) => { console.error(error); process.exit(1); });
657+
EOF
658+
659+
# The TypeScript SDK's framework adapters, against REAL framework releases —
660+
# the counterpart of `failproofai-sdk-integrations` above. `failproofai-ts-sdk`
661+
# runs the adapters with no framework installed, which proves their logic and
662+
# nothing about whether it ever reaches a framework: the first release passed
663+
# 242 unit tests with every adapter recording nothing in an ES-module app.
664+
#
665+
# Each fixture under `sdk/typescript/integration/fixtures/` is a consumer
666+
# project with its own lockfile pinning one framework release. The packed
667+
# tarball is extracted into each, and one agent is run as BOTH an ES module and
668+
# CommonJS, because the two module systems load different copies of a
669+
# dual-published framework. A fixture that fails to install fails the job —
670+
# there is no skip path to read as green.
671+
failproofai-ts-sdk-integrations:
672+
name: failproofai-ts-sdk-integrations (${{ matrix.shard }}, node ${{ matrix.node-version }})
673+
runs-on: ubuntu-latest
674+
timeout-minutes: 30
675+
defaults:
676+
run:
677+
working-directory: sdk/typescript
678+
strategy:
679+
fail-fast: false
680+
matrix:
681+
# Sharded by what a shard needs installed, because the whole suite —
682+
# four frameworks at both ends of their ranges, Bun and Deno parity over
683+
# every fixture, and five real `next build`s — is ~23 CPU-minutes, too
684+
# much for one runner inside a timeout. Each shard installs only its own
685+
# fixtures (FAILPROOFAI_IT_FIXTURES) and runs only its own files.
686+
#
687+
# frameworks: Node's oldest and newest supported majors — the ESM/CJS
688+
# split this job exists for behaves differently once `require(esm)` is
689+
# unflagged. runtimes and nextjs: one Node, because what they vary is
690+
# the runtime or the bundler, not Node.
691+
include:
692+
- shard: frameworks
693+
node-version: "20.x"
694+
fixtures: ai-4,ai-5,ai-6,ai-7,langchain-0.3,langchain-1,langchain-dup-core,mastra-0,mastra-1,llamaindex-0.11,llamaindex-0.12,types,vanilla
695+
files: integration/ai.test.ts integration/langchain.test.ts integration/mastra.test.ts integration/mastra-coverage.test.ts integration/llamaindex.test.ts integration/types.test.ts integration/vanilla.test.ts
696+
- shard: frameworks
697+
node-version: "24.x"
698+
fixtures: ai-4,ai-5,ai-6,ai-7,langchain-0.3,langchain-1,langchain-dup-core,mastra-0,mastra-1,llamaindex-0.11,llamaindex-0.12,types,vanilla
699+
files: integration/ai.test.ts integration/langchain.test.ts integration/mastra.test.ts integration/mastra-coverage.test.ts integration/llamaindex.test.ts integration/types.test.ts integration/vanilla.test.ts
700+
- shard: runtimes
701+
node-version: "24.x"
702+
fixtures: ai-4,ai-5,ai-6,ai-7,langchain-0.3,langchain-1,mastra-0,mastra-1,llamaindex-0.11,llamaindex-0.12,runtimes
703+
files: integration/runtimes.core.test.ts integration/runtimes.bun.test.ts integration/runtimes.deno.test.ts
704+
- shard: nextjs
705+
node-version: "24.x"
706+
fixtures: nextjs,langchain-1,ai-7,mastra-1,llamaindex-0.12
707+
files: integration/nextjs.test.ts
708+
steps:
709+
- uses: actions/checkout@v7.0.1
710+
711+
- uses: actions/setup-node@v5
712+
with:
713+
node-version: ${{ matrix.node-version }}
714+
cache: npm
715+
cache-dependency-path: |
716+
sdk/typescript/package-lock.json
717+
sdk/typescript/integration/fixtures/*/package-lock.json
718+
719+
- name: Install dependencies
720+
uses: nick-fields/retry@v4
721+
with:
722+
max_attempts: 3
723+
timeout_minutes: 5
724+
command: cd sdk/typescript && npm ci --no-audit --no-fund
725+
726+
# `test:integration` builds, packs, `npm ci`s this shard's fixtures and
727+
# runs its files; the fixture installs are the network-heavy part, so
728+
# retry them as a whole rather than failing the job on a registry blip.
729+
- name: Test against real releases (${{ matrix.shard }})
730+
uses: nick-fields/retry@v4
731+
env:
732+
FAILPROOFAI_IT_FIXTURES: ${{ matrix.fixtures }}
733+
with:
734+
max_attempts: 2
735+
timeout_minutes: 25
736+
command: cd sdk/typescript && npm run test:integration -- ${{ matrix.files }}
737+
487738
test:
488739
runs-on: ubuntu-latest
489740
# The retry above nominally allows 3 attempts x 10 minutes. Capping the job

‎.github/workflows/osv-scanner.yml‎

Lines changed: 2 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -62,7 +62,7 @@ jobs:
6262
# nothing at all — not this job, which was only ever given `bun.lock`, and
6363
# not Dependabot, which had no `cargo` ecosystem — for a TLS stack that
6464
# compiles into a root-installed system service.
65-
- name: Scan bun.lock and Cargo.lock for known-vulnerable / malicious dependencies
65+
- name: Scan every lockfile for known-vulnerable / malicious dependencies
6666
id: scan
6767
uses: google/osv-scanner-action/osv-scanner-action@f4cfcc01edc9c8b756a9b873b7a623ca674da51e # v2.3.8
6868
with:
@@ -79,6 +79,7 @@ jobs:
7979
--lockfile=Cargo.lock
8080
--lockfile=fp-cloud-cli/uv.lock
8181
--lockfile=sdk/python/uv.lock
82+
--lockfile=sdk/typescript/package-lock.json
8283
# Only the schedule run notifies — nothing on main touched the lockfile,
8384
# so nobody is watching it the way a PR author watches their own checks
8485
# or a push failure shows up against the commit they just merged. Same

0 commit comments

Comments
 (0)