Commit f2aa5b4
[sdk/typescript] Add @failproofai/sdk, the TypeScript telemetry SDK (#830)
* feat(sdk): add @failproofai/sdk, the TypeScript telemetry SDK
The counterpart to the Python SDK at sdk/python: the same 15 events, the
same wire format, the same spool directory, the same Evaluator v2
protocol. A fleet running Node agents and Python agents now writes into
one pipe, and the dashboard cannot tell which wrote what.
Three surfaces, mirroring the Python ones:
* Scopes — session(), agent(), toolCall(), with identity carried on
AsyncLocalStorage. Each has a callback form and a `using`-compatible
.open(). A synchronous body stays synchronous; wrapping every call in
a promise would break the one case that genuinely cannot await.
* Adapters — instrument() wires LangChain.js/LangGraph.js, the Vercel AI
SDK, Mastra and LlamaIndex.TS.
* event.* — the 15 methods, with the same validation. The promoted
columns are checked at the boundary because ingest answers 200 OK and
stores NULL for a value it cannot read, so the alternative is a
dashboard that is quietly missing rows.
Two places the port deliberately diverges, both because the language
differs rather than because the design does:
* The Vercel AI SDK exports functions from an ES module, and an ES
module namespace is immutable — there is nowhere to patch. It is
served by the two extension points that SDK documents: an
OpenTelemetry-shaped tracer for experimental_telemetry, and a
LanguageModelV2Middleware. Using both records each call once.
* Managed evaluator source is PARSED AND INTERPRETED, not eval'd and not
handed to node:vm. Python's AST allowlist plus eval does not transfer:
x["constructor"]["constructor"]("…")() reaches arbitrary code through a
key computed at runtime, which no source-level check can see, and a vm
context has its own Function. Every property read goes through one
function that checks the actual key at the moment of the read. The
worker_threads sandbox around it — V8 heap limits, a wall-clock
terminate(), a bounded result, a concurrency cap — is the RESOURCE
bound, and it fails closed: no sandbox means refusal.
Zero runtime dependencies, enforced by a test and by the build. Dual
ESM + CommonJS, Node >= 20.9. 242 tests in the package, plus an 18-case
pipeline test in __tests__/ci/.
Infrastructure: a CI job across four Node majors that smoke-tests the
packed TARBALL (both module systems, --omit=peer, events read back off
disk, and the evaluator sandbox resolved through the package's own
export — the one part that cannot be exercised from inside this repo); a
release workflow with the same preflight/build/publish/verify/bump shape
as its Python sibling; the new lockfile added to the OSV scan; and
sdk/typescript excluded from the root tsconfig and eslint config, so this
project's dependency tree cannot decide whether that package's
zero-dependency claim holds.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BJ3C2FkvrjzkxNUqGxvgSW
* docs: add the TypeScript SDK reference page and cross-link it
The Python custom-agents reference is the page a customer lands on from
PyPI, and it was the only place either SDK was documented. A TypeScript
reference sits beside it now, registered in the English navigation; the
translate pipeline picks up the other fourteen locales on its next run,
which is what it is for.
The cross-link runs both ways, and both sides say the same thing in the
same place: the two SDKs write the same events into the same spool, so a
fleet with Node agents and Python agents produces one set of sessions,
not two. That is the fact a reader needs before they start choosing, and
it does not appear anywhere the choice is actually made otherwise.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BJ3C2FkvrjzkxNUqGxvgSW
* fix(sdk): stop the TS SDK's tests inheriting the monorepo root's tooling
Two failures the isolated CI job found and a local run structurally
cannot, because both only appear once this package stands on its own.
**Vite read the ROOT's postcss.config.mjs.** It searches upward from the
project root for a PostCSS config and finds the one at the repository
root, which requires `@tailwindcss/postcss` — a root devDependency that
is deliberately absent from this package's node_modules. Vite then dies
before a single test runs. Locally the root's node_modules is present
and Node's resolution walks up into it, so the search succeeds and
nothing looks wrong; in CI, where the isolation is real, it is fatal.
An inline empty `css.postcss` turns the search off. This package has no
CSS at all, so the only thing that search can do here is find somebody
else's tooling. Verified by hiding `@tailwindcss/postcss` from the root
node_modules and re-running: 242 passed.
**vitest 2.1.9 carried seven fixable advisories.** The Supply Chain gate
flagged vite 5.4.21 and esbuild 0.21.5 underneath it, one Critical and
one High. osv-scanner.toml says to prefer fixing over ignoring, so this
moves to vitest 5 — the major the root project already uses — which
pulls vite 8.3.0 and drops the vulnerable esbuild entirely. The Vitest 4
pool rework replaced `poolOptions.forks.singleFork` with the top-level
`fileParallelism`, which says the same thing more plainly: run files one
at a time, so a test never observes another file's patched prototype or
competes with itself for the sandbox semaphore.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BJ3C2FkvrjzkxNUqGxvgSW
* ci: prove the TS SDK's Node floor with the artifact, not the test runner
The 20.9 leg failed to start: vitest 5 pulls vite 8, which pulls
rolldown, which imports `styleText` from `node:util` — added in Node
20.12. Nothing about the package needs it; 20.x, 22.x and 24.x all
passed.
So the leg says what is actually true. The TEST RUNNER's floor is not
the PACKAGE's floor, and the honest way to hold the floor at 20.9 is to
prove it with the thing a consumer receives: the floor leg builds, packs,
installs the tarball and runs it, and skips the suite the runner cannot
start there.
The alternative — pinning vitest back to something a 2023-era Node can
load — is the tail wagging the dog. It means carrying the CVEs vitest 5
fixed (one Critical, one High, flagged by the Supply Chain gate on the
previous push) so that a test runner can start on a release nobody runs
the tests on.
`__tests__/ci/ts-sdk-pipeline.test.ts` asserts both halves: the floor is
in the matrix, at least one leg runs the suite, Build/Pack/smoke are
ungated, and exactly Typecheck/Lint/Test carry the gate. A later edit
that quietly ungates the suite, or drops the floor leg's smoke test,
fails there rather than in a release.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BJ3C2FkvrjzkxNUqGxvgSW
* test(sdk/typescript): add a real-framework integration suite; patch the copy the app loads
The adapters were only ever tested against src/ with no framework installed,
so nothing proved an adapter reached a real framework. They didn't: every
adapter resolved its framework with createRequire, which names the CommonJS
build of a dual-published package, and patched that copy. An ES-module app
(the default for new TypeScript projects) loads the other copy, so
instrument() reported success and recorded nothing.
integration/ installs real framework releases from per-fixture lockfiles,
extracts the PACKED tarball into each, and runs one agent.ts as both ESM and
CJS. Expected traces are the Python SDK's golden output for the same program.
compat.requireModuleCopies() now returns the copy the application's imports
reach (by the entry point's module system, resolved through the package's
exports map with ESM conditions) plus the CommonJS copy if something already
required it, and never loads a copy speculatively. The LangChain adapter now
patches every copy; its mapping is still the pre-parity one.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* ci(sdk/typescript): run the framework adapters against real releases; test copy selection
A `failproofai-ts-sdk-integrations` job, the counterpart of the Python SDK's
`failproofai-sdk-integrations`: packs the SDK, installs every fixture from its
lockfile, and runs each agent as ESM and CJS on Node 20 and 24.
test/copies.test.ts pins which copy of a dual-published package
requireModuleCopies() returns, from real ESM and CJS entry points against a
fake dual package on disk: the ESM build for an ES-module app, the CJS build for
a CommonJS app, both when the CJS copy was already required, and never a
speculatively loaded second copy.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): ship CommonJS declarations and node10 subpath types
The exports map handed the ESM .d.ts files to `require`, so TypeScript read
every CommonJS import of the package as an ES module: TS1479 on every import
under `module: node16` (and `nodenext` before TS 5.8), and attw
"Masquerading as ESM" for node16-from-CJS on every entry plus
node16-from-ESM on ./sandbox-worker. `moduleResolution: node` (node10, what
`module: commonjs` implies) resolved the root and no subpath at all.
- the CJS build now emits declarations; under dist/cjs's
`"type": "commonjs"` they are CommonJS declarations
- exports use nested import/require conditions, each with its own `types`;
./sandbox-worker (CommonJS only) points its types at dist/cjs
- `types` points at the CJS declarations and `typesVersions` maps each
subpath for node10; exports-aware resolvers ignore it
- finalize-build walks nested conditions, refuses a half whose types and
JavaScript live in different directories, and checks typesVersions
integration/types.test.ts typechecks a consumer of every entry point under
six consumer tsconfigs on TS 5.9.3 and 5.4.5 (the minimum: the oldest with
`module: preserve`), skipLibCheck off, and runs attw over the packed tarball.
test/packaging.test.ts learns the nested condition shape and asserts it.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): make the Vercel AI SDK adapter record real ai 4-7 traces
Verified against real `ai` releases, the adapter was broken in six ways:
instrument("ai") always threw (setGlobalTracerProvider called unbound) and
its "don't overwrite a provider" check read the always-present proxy;
startActiveSpan auto-ended spans, so streamText lost the half the SDK ends
itself (endWhenDone:false); the README call sites did not typecheck on ai
5/6; ai 7 was unsupported, silent, and ERESOLVE'd on install; v3 usage
objects and {unified,raw} finish reasons lost tokens and wrote objects into
stop_reason; and a bare wrapModel call was dropped or pinned to a phantom
"main" agent.
The adapter now follows the Python mapping rule: one operation = one agent
named by functionId (never an id), each model step a request/response pair
with integer tokens and a string stop reason, each tool a pair on the
model's toolCallId, a failure recorded once where it happened. v4-v6 attach
through a structural OTel tracer that never ends spans for the caller; v7
through its Telemetry integration interface (per call via telemetry(), or
globalThis.AI_SDK_TELEMETRY_INTEGRATIONS via instrument()). telemetry()
returns both, and v6's callId-less integration events are ignored, so one
call site works on every major and nothing records twice. A bare wrapModel
call becomes its own run named after the model unless an agent() owns it.
Public types are structural and typecheck against real ai 4/5/6/7 types.
Peer range widened to ai >=4.0.0 <8. integration/ adds ai-4..ai-7 fixtures
(pinned lockfiles, ESM + CJS, a nodenext typecheck of the README call sites).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): make the LangChain adapter draw the Python adapter's trace
The LangChain.js / LangGraph.js adapter recorded every LangGraph node as a
nested agent, every tool call under its run id, a run-level `error` event per
layer a failure unwound through, and nothing at all for sessions keyed by
metadata. It is now a port of sdk/python's integrations/langchain.py, held to
that adapter's golden output by the real-framework suite:
- root run = agent; a LangGraph node = hook_triggered/hook_completed with
trigger_event "graph_node" (Python's `_node_of` rules); a compiled subgraph
= nested agent "root/node"; machinery emits nothing unless `includeChains`
- tool_use/tool_result carry the MODEL's tool_call_id (core 1.x passes it;
on core 0.3 it is recovered from the ancestor's tool_calls)
- failure is carried by model_response.error, hook_completed failed and
agent_end failed; a standalone `error` only when no span owned it
- interrupt() -> human_wait + agent_pause, Command resume -> agent_resume +
human_input on the same agent, including a resume taken by ANOTHER process
(the pause id rebuilt from the checkpoint namespace, as LangGraph derives it)
- session order: sessionId option, metadata.failproofai_sdk_session_id,
ambient scope, session_id/conversation_id/thread_id, root run id
- streaming folded into model_response (fw_streamed, fw_chunks, fw_ttft_ms);
.batch() roots stay separate; aborts end `cancelled`, including the root
LangGraph.js 1.x abandons without an end callback
- langchainHandler() works without instrument(), never double-records with
it, and uninstrument() closes open runs as cancelled
- the handler is awaited (awaitHandlers) so events keep the caller's async
context, and configure() receives it as an input handler so the call's
metadata is not dropped
Adds a langchain-0.3 fixture (@langchain/core 0.3.80, @langchain/langgraph
0.4.10) run over the same expectations, so the declared >=0.3.0 <2 range is
tested at both ends.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): make the Mastra adapter record real Mastra runs
Verified against @mastra/core 1.68.0 and 0.24.9, the adapter recorded
nothing in an ES-module app, never saw a tool, labelled every agent
"mastra-agent", collapsed a multi-step loop into one model pair with no
model or tokens, ended streams before they ran, and dropped a bare
wrapTool() call. Rewritten on patch points every run goes through, on
every copy of @mastra/core the app loads (compat.requireModuleCopies):
- Agent.generate/stream (+ legacy/VNext): the agent span, named after
the agent, nested under an enclosing scope or the delegating agent.
stream() ends on Mastra's onFinish/onError/onAbort, not on return.
- Agent.resolveModelConfig: the resolved model goes back behind a proxy
observing doGenerate/doStream, so each LLM step is its own
model_request/model_response with model id, tokens, stop reason.
- Agent.convertTools: wraps the loop's converted tools, so every tool
(including ones built before instrument()) records with the model's
toolCallId, attributed to the agent whose loop called it.
- Run._start/_resume and DefaultExecutionEngine.executeStep: workflow
runs as agents, steps as workflow_step hooks; Mastra's own internal
workflows (the agentic loop) are skipped.
Mastra 1.x's ObservabilityExporter was considered and rejected: it only
fires for agents registered on a Mastra instance with
@mastra/observability configured, so a bare Agent records nothing.
Peer floor raised 0.10.0 -> 0.20.0, where resolveModelConfig and
Run._start first exist; earlier lines route model calls through the AI
SDK and would record no steps.
Adds integration/mastra.test.ts with mastra-1 (1.68.0) and mastra-0
(0.24.9) fixtures, ESM and CJS, and unit tests for the readers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): make the LlamaIndex.TS adapter record real runs
The adapter recorded nothing in an ES-module app, split every legacy
LLMAgent run into two sessions with an agent left open, dropped streamed
token usage and every model name, and dropped bare llm.chat() calls. Nothing
recorded an agent().run() at all: @llamaindex/workflow emits no run or step
events on the callback bus.
It now attaches in two places. The callback bus (@llamaindex/core/global,
subscribed on each copy the app uses, via compat.requireModuleCopies) carries
model, tool, retrieval, query and legacy-runner events. AgentWorkflow's
runStream is patched to open the run and attach to the context's
__internal__call_context / __internal__call_send_event middleware hooks, the
surface workflow-core's own withTraceEvents/withState middleware uses, so
steps become hooks and the stop event ends the run. Bus events are placed by
an AsyncLocalStorage frame bound around each step, or by LlamaIndex's own
EventCaller chain (legacy runners, query engines, providers' chat).
Same tree as the Python adapter: run = agent named after the agent,
multi-agent handoff = nested agents, step = workflow_step hook, bare call =
its own run. Streamed usage is read off the chunk that carries it; the model
name comes from the caller chain or the agent the step names.
The floor moves from 0.9.0 to 0.11.4, the first llamaindex on the workflow
1.1 runtime; 0.9-0.11.3 ship workflow 1.0, whose AgentWorkflow has no
runStream. Integration fixtures pin 0.11.4 (the floor) and 0.12.1.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(sdk/typescript): state the tested framework ranges and the Python mapping
The adapter table claimed LangGraph nodes become agent spans, Mastra's
createTool is patched and LlamaIndex is never patched — none true any more —
and stated no supported versions at all. It now names the range each adapter
is tested against, the mapping it shares with the Python SDK, how the ESM/CJS
dual-package case is handled and where it cannot be (bundled output), the
ai 7 `telemetry:` spelling, and langchainHandler() without instrument().
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): bound what a long-running process retains; find the app from its entry
From an adversarial review, each reproduced first:
- RunTracker.links was never pruned when a run ended, so it filled to its
10k FIFO cap with finished runs and then evicted the links of LIVE ones: a
LangChain model call that outlived ~10k other runs lost its model_response to
"could not resolve a session". Links now go when their run closes (emit()
drops them after a closing event, endAgent after agent_end, LangChain after
every end path), and a re-link refreshes an entry's age.
- A LangGraph run paused on a human and resumed by another worker was held
open here forever, taking a tracker slot live runs need, and every callback
scanned all open agents. isOpen() is O(1); paused sessions are forgotten
(never closed — the other worker closes them) after PAUSED_SESSION_TTL_MS or
when they fall off the session cap. A late resume takes the cross-worker path.
- Framework resolution was anchored at process.cwd() alone: a service started
from / found no framework, and in a monorepo it patched the hoisted root copy
instead of the app's own. It now resolves from the entry script's directory
and the working directory, preferring the entry unless it is a launcher inside
node_modules, and a copy already required wins.
- `@internal` test hooks no longer ship in the published declarations
(stripInternal).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): stop the Mastra adapter recording after uninstrument()
Model proxies and tool wrappers built while instrumented called
ensureTracker(), which lazily built a FRESH tracker after teardown. A
stream in flight at uninstrument() therefore kept emitting model/tool
events - attributed to a phantom "main" agent inside any open scope,
plus a second agent_end - and anything an instrumented run had built
stayed live forever.
- One tracker per installation, created by install() and dropped by
uninstall(); an `enabled` flag checked on every recording path. Every
proxy, wrapper, span and frame remembers the tracker it began on and
becomes a pass-through once that tracker is gone, including after a
later instrument().
- uninstall() switches off first, then closes what is open: model steps
(stop_reason "cancelled"), tool calls and workflow steps (all marked
fw_incomplete), then agents "cancelled".
- Open model/tool/step spans are tracked in bounded registries that every
end path prunes (success, error, abort, cancel); a never-consumed
stream is held only until teardown closes it.
- A model proxy or tool wrapper from an earlier run is re-wrapped for the
current run instead of being skipped as already wrapped.
- instrument()'s captureLimit is honoured even when wrapTool() recorded
first (it used to be ignored). wrapTool() keeps working without
instrument() on its own self-contained tracker.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): Mastra calls cut off by uninstrument() stop as "incomplete", like LangChain's
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): keep concurrent LlamaIndex runs on one shared object apart
Runs were keyed by the object that owns them (the first computedCaller), and
a query engine or LLMAgent built once and shared by every request is the same
object for every call: request B's query-start found A's run, treated itself
as nested and recorded nothing; B's retrieval landed in A's session and B's
answer was never recorded. Legacy chats nested B under A and sent A's later
model calls to B.
Query and legacy runs are now keyed by LlamaIndex's own EventCaller - a fresh
object per @wrapEventCaller invocation, bound in LlamaIndex's AsyncLocalStorage
and chained through .parent - and an event belongs to a run only when that
exact invocation is on its chain. This is used instead of prototype-patching
query/chat because @wrapEventCaller binds the method onto each INSTANCE at
construction, so a prototype patch would miss every engine built before
instrument(). A bus with no EventCaller falls back to owner matching, where a
run owned by the object starting a new run is never taken as its parent.
Also:
- attach() throws while an install is live instead of orphaning its bus
subscriptions and runStream patch; documented @internal.
- Every tracker link a run makes (run key, leaf keys, workflow step keys) is
unlinked on every end path, so completed runs leave no residue and the
tracker's FIFO cap never evicts a live run's link. Adds RunTracker.unlink().
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): stop instrument("ai") taking the global OTel slot; close cancelled/errored streams
instrument("ai") on ai 4-6 registered FailproofTracer as the process-wide
OpenTelemetry tracer provider whenever the slot was empty. OpenTelemetry
refuses every later registration, so a customer's NodeSDK.start() further
into startup was silently refused and their http/pg/Next.js spans went to a
tracer that exports nothing. Taking the slot is now opt-in
(instrument("ai", { registerGlobalTracer: true })); by default instrument("ai")
on 4-6 logs one warning naming the call-site telemetry()/wrapModel() paths,
and registerGlobalTracer: false silences it. The v7 global integration list
is additive and per-call integrations replace it, so it is kept.
middleware()/wrapModel() observed streams with pipeThrough(TransformStream),
whose flush never runs on a consumer cancel or a stream error: the model call
stayed open forever and a standalone ai-model run never got agent_end. It now
uses a pull-based re-stream (core.observeStream, extracted from mastra.ts,
which now calls it): cancel closes the pair with stop_reason "cancelled" and
ends the run cancelled (the cancel reaches the provider stream); an error
closes it with the error and ends the run failed.
Every tracer span, v7 call/model/tool and middleware request now unlinks its
RunTracker parent link when its last event is out, including tools a v7
operation abandoned. Adds RunTracker.unlink(). 20k completed operations
leave no residue and no longer evict a live run's links.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(sdk/typescript): changelog for the long-running-server fixes and the instrument("ai") opt-in
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): record LangChain roots started through a nested @langchain/core; cover every common LangChain.js surface
Coverage sweep of the LangChain.js adapter against real releases, both module
systems, both ends of the peer range:
- langchain-0.3 + langchain-1: LCEL (prompt|model|parser), RunnableSequence,
RunnableParallel, RunnableLambda, a custom retriever, a VectorStore
.asRetriever(), a RAG chain, .batch() x3, .streamEvents() v2, chain
.stream(), withStructuredOutput(zod), bindTools with tool_choice — each
under instrument() AND through an explicit langchainHandler(). Expected
traces are the Python adapter's for the equivalent program.
- langchain-1: the v1 `langchain` package's createAgent (1.5.12), plain,
streamed, explicit-handler and with middleware (agent-v1.ts).
- langchain-dup-core (new fixture): an app on core 1.2.12 whose vendored
provider pins core 0.3.80, so npm nests a second copy under it.
Bugs found and fixed:
- Runs a nested @langchain/core copy started as ROOTS (a provider's model,
tool, retriever or runnable invoked directly) were recorded nowhere under
instrument(): only the app's copy of CallbackManager was patched.
LangChain has no cross-copy hook (registerConfigureHook keys on a
module-private Symbol), so install now finds nested copies on disk
(node-require.nestedCopies, incl. the pnpm store) and patches each
in-range copy's reachable build(s).
- createAgent's model node returns a LangGraph Command; the payload view
handed it to truncate() whole, so its messages were captured as
LangChain's {lc, type: "constructor", id, kwargs} envelope.
Harness: additive — transpile every agent*.ts program of a fixture, and
runAgent takes an optional program name.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): AI SDK coverage sweep — embeddings, abandoned streams, tool-call content
Walk every commonly used ai 4–7 surface against the packed SDK (new
surfaces.ts per fixture, ESM + CJS): agent classes, embed/embedMany,
structured output, tool features, stream consumption styles, reasoning
models and concurrency. Fixes what that exposed:
- embed/embedMany inside an agent() scope or a tool are model calls of the
enclosing agent, not nested ai.embed agents (bare calls stay their own run)
- v4–v6: an aborted stream closes its open model/tool spans as cancelled and
ends the agent cancelled (was: success with an unpaired model_request)
- v4–v6: a stream the SDK never ends (client disconnect, never read, v4
mid-stream provider error) is closed when its root span is collected
(FinalizationRegistry, fw_abandoned) instead of staying open forever
- v4–v6: tool calls in model_response content use one shape with parsed
input (was v4 {toolCallType,args:"<json>"} / v5-6 input as a JSON string)
- v7: an aborted model call closes as cancelled (was "incomplete") and a
failed/aborted one keeps its model id
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): Mastra coverage sweep — memory threads, networks, suspend/resume, tripwires
Integration cases (ESM + CJS, @mastra/core 0.24.9 and 1.68.0) for every
commonly used Mastra surface: agents/workflows fetched from a Mastra
instance, 10 concurrent runs of one Agent, @mastra/memory threads, agent
networks, agent-in-tool delegation, branch/parallel/dowhile/foreach/nested
workflows, createStep(agent), run.stream(), suspend/resume, input/output
processors and tripwires, MCP tools over a local stdio server, structured
output (incl. a second structuring model), maxSteps with a failing tool, and
streamed usage from an OpenAI-compatible endpoint.
Bugs they exposed, fixed in the adapter:
- a memory thread was not the session: every turn of one conversation was a
new session, and 1.x's memory: { thread } never reached fw_thread_id;
- agent.network() shattered into one root session per routing decision
("Routing Agent" x2 + the delegate), and on 0.x recorded Mastra's own
network workflows plus the delegate's internal loop steps as hooks;
- workflow suspend/resume was two unrelated sessions with no HITL events;
now human_wait + agent_pause ... agent_resume + human_input on one span,
with deterministic ids and a cross-process resume that closes them;
- a stream() blocked by a processor tripwire left its agent open forever.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): LlamaIndex coverage sweep — chat engines, retriever names, plain workflows
Real-framework cases (llamaindex 0.11.4 + 0.12.1, ESM + CJS) for every
commonly used LlamaIndex.TS surface, checked against the Python adapter's
golden traces. Three bugs they exposed:
- A chat engine (Simple/Context/CondenseQuestion) dispatches nothing of its
own, so its retrieval and model call were separate root runs in separate
sessions, named after the model. Its run boundary is now read from
LlamaIndex's own EventCaller storage (an own `run` on that one object):
a top-level invocation is one run named after its class, ended when it
returns. The same signal closes a model call, query or legacy task that
THREW, which previously stayed open until the reaper.
- Retrievals were named "retriever"; Python names them after the retriever
class. BaseRetriever.prototype.retrieve now binds the retriever for its
retrieve-start.
- createWorkflow() workflows recorded no run and no steps. On workflow-core
>=1.1 its exported AsyncContext.Variable sees every step handler; a
context's burst of steps is now a "Workflow" run with hook pairs.
Also covered: multiAgent() with 3 agents, responseFormat, FunctionTool.from,
QueryEngineTool, parallel tool calls, a mid-stream provider error (run closes
failed despite LlamaIndex's unhandled rejection), and 10-way concurrency on a
shared agent, query engine and chat engine. Unit tests for each fix and for
@llamaindex/openai's include_usage stream chunk.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): flush() waits for events emitted during an in-flight flush
flushNow() — what the public flush() awaits — handed back a flush that was
already running, which had drained the queue before any event emitted since.
So `await failproofai.flush()` could resolve with the newest events still in
memory: exactly the emit-flush-exit case it exists for. It showed up as a
~1-in-6 flaky unit test. It now flushes twice: the second pass drains what the
first could not see, or joins a newer flush that already did.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(sdk/typescript): run every framework fixture under Bun and Deno; fix Deno ESM apps recording nothing
Deno sets require.main to null (not undefined) for an ES-module entry, so
entryIsCommonJs() read every ES-module Deno app as CommonJS and instrument()
patched the frameworks' CommonJS copies while the app ran the ES-module ones:
LangChain, Mastra and LlamaIndex reported success and recorded nothing.
Adds a pinned runtimes fixture (bun 1.4.2, deno 2.9.6 from npm, located by the
harness), bun-*/deno-* formats, a Node-parity suite over all 184 framework
scenarios per runtime, a core smoke of every scope and event method, the
short-lived-handler (Lambda) cases, and Deno npm: specifiers.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): instrument() records under Next.js serverExternalPackages; importing the SDK no longer breaks an Edge build
Next.js: next start is a CommonJS launcher, but Next loads every server-external
package with import() (Turbopack and webpack alike), so the app runs the
frameworks' ES-module copies. The SDK read the launcher as a CommonJS app and
patched the CommonJS copies: instrument() in instrumentation.ts reported success
for LangChain, Mastra and LlamaIndex and recorded nothing. Inside a Next server
(NEXT_RUNTIME set) the ES-module copy is now the one patched
(appImportsReachCommonJs in node-require.ts, used by requireModuleCopies).
Edge: an Edge-runtime route importing the SDK failed next build outright
("Native module not found: node:fs"). The package now maps the edge-light,
workerd, worker and browser conditions to a no-op build (src/edge/) with the
same public surface that logs one line on first use; types stay the real ones.
Adds a Next.js 16 App Router fixture (one route per framework, a server action,
instrumentation.ts, an Edge route) built four ways - Turbopack/webpack x
default(bundled)/serverExternalPackages - with each route's trace compared to
the same program's trace under plain Node. Drops the short-lived-handler cases
from the runtimes suite (not a deployment target).
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(sdk/typescript): reconcile the parallel coverage sweeps
Each sweep was green alone; merged, three things disagreed:
- the edge no-op build's parity test did not know two helpers the ai and
llamaindex sweeps export for their unit tests (toolCallsOf,
eventCallerStorage) — both marked @internal and listed as such;
- runtime parity compared event ORDER within a session for the Mastra sweep's
`concurrent` scenario, whose ten runs share one session and interleave by
scheduling on every runtime — such scenarios now compare the multiset;
Mastra 0.x `tripwire-output`, a documented limit its own suite skips, is
skipped here too;
- Deno's node:http keeps idle keep-alive sockets open through close(), so the
loopback provider in `usage-openai-compatible` held the process past its
timeout — the fixture now closes them.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(sdk/typescript): withFailproofai() for Next.js, and a warning when Next bundles a framework
`next build` bundles LangChain, Mastra and LlamaIndex into its server output by
default, so `instrument()` patched the node_modules copy the app never runs,
reported success, and recorded nothing.
`@failproofai/sdk/next` exports `withFailproofai(nextConfig)`, which adds the
packages the adapters need (and the SDK itself, so instrumentation.ts and the
routes share one copy) to `serverExternalPackages`, keeping the app's own list,
any config form (object, sync or async function), and leaving out anything the
app lists in `transpilePackages`.
`instrument()` on a Next.js Node server now warns once per adapter that cannot
reach a bundled copy — LangChain, Mastra, LlamaIndex, never the AI SDK — unless
the app is configured: the list the wrapper records when Next evaluates the
config, a standalone server's resolved config, or a hand-set
FAILPROOFAI_NEXT_EXTERNALS=1.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(sdk/typescript): build each Next.js variant from scratch
`npm pack` stamps every file with one fixed mtime, so when the harness put a
newly packed SDK into the fixture, webpack's persistent cache (kept in the
distDir) saw it as unchanged and bundled the PREVIOUS run's SDK. The webpack
variants were silently testing stale SDK code; Turbopack hashes content and was
unaffected. Found because the webpack build printed none of the new Next.js
warnings while Turbopack printed all three.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* ci(sdk/typescript): shard the real-framework suite; guard the shard lists
Frameworks at both ends of their ranges, Bun and Deno parity over every fixture,
and five real `next build`s are ~23 CPU-minutes — too much for one runner inside
a timeout. The job is now a matrix: frameworks (Node 20 and 24), runtimes, and
nextjs, each installing only its own fixtures and running only its own files.
The shard lists are hand-maintained, so a new integration file left out of every
shard would never run in CI. __tests__/ci/ts-sdk-integration-shards.test.ts
fails when a file is uncovered or a shard names a file or fixture that does not
exist.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* docs(sdk/typescript): Next.js setup, runtimes, per-framework notes and streamed-usage caveats
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* test(sdk/typescript): cover the ./next entry point in the declarations suite
The attw check pins the public entry-point list, so the new `./next` export
failed it. It is now listed, and the consumer file every resolution mode
typechecks imports withFailproofai and proves its result is typed.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* feat(sdk/typescript): instrument your own agent — guide and a CI-proven example
For a client whose agents are hand-built TypeScript with no framework. The Python
SDK's answer carries over one to one: every such agent already has three places
(where a run starts and ends, the one function that calls the model, the one
function that runs tools), and the scopes at those three are the whole
integration, producing the same trace the adapters do.
examples/research-agent.ts is the TypeScript twin of the Python SDK's
research_agent.py — a real OpenAI tool loop instrumented by hand, pairing model
calls on requestId, timing them, closing them on failure, and reusing the
model's tool-call ids. The `vanilla` fixture IS that file (a test holds them
byte-identical) and runs it as ESM and CJS on the real openai client against a
local OpenAI-compatible server: parallel tools, a failing tool, a failing model.
README and docs gain a "your own agent — no framework" section.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* fix(sdk/typescript): run CommonJS evaluator files; name the real ai import in warnings
failproofai-evaluator is the ESM build, and a CommonJS evals file builds its
Evaluator from dist/cjs, a second copy of the class, so the loader's instanceof
check refused it. It now reads a Symbol.for brand both builds stamp. A
packaging test runs a CommonJS and an ESM file through the built bin.
The ai adapter's warnings told readers to call failproofai.ai.telemetry(),
which the root package does not export; they now name the @failproofai/sdk/ai
import.
* docs(skill): failproofai-sdk covers TypeScript agents and the evaluator worker
New references/typescript.md (install, names, adapters, Next.js and bundlers,
the three-edit-site wiring for a hand-built loop, shutdown, verification) and
references/evaluator.md (writing, running and deploying the eval pod in Python
and TypeScript). The description no longer routes to the retired
agenteye-evaluator, and agents/openai.yaml takes the policy: nesting the skills
repo already carries so the next sync does not revert it.
test/skill-snippets.test.ts parses every TypeScript block in the skill and
fails when one calls an SDK name that no longer exists.
* fix(sdk/typescript): what six user-style runs against Cloud found
- Timestamps: events in one millisecond tied, and the dashboard showed a
tool_result before its tool_use. The digits below the millisecond are now
a per-process sequence (clock.ts), re-anchoring if the wall clock steps back.
- Exit: SIGTERM left runs rendered as running forever. The exit hook closes
open tools (ProcessExit error) and agents (error + agent_end failed),
including adapter-opened runs; a run paused on a human stays open for the
process that resumes it.
- configure() reached only its own copy of the SDK, so a Next.js route bundled
without withFailproofai reported environment dev. environment and baseDir
are process-wide now (shared.ts).
- An agent() wrapper around a framework agent of the same name joins it
instead of opening a self-parented duplicate.
- LlamaIndex 0.12 tool errors no longer read 'Error: Error(Error): ...';
error_type is the class for errors whose name is just 'Error'; an
unawaited instrument() is warned about by the LangChain adapter.
- The vanilla example records the model's tool calls, keeps its error path to
the provider call, and survives malformed tool arguments.
The skill's TypeScript page covers naming, owning the session id, adapter
options, streamed-usage flags, verifying under a running daemon, Next.js
configure() placement and the new shutdown behaviour.
* fix(sdk/typescript): close every open leaf at exit; a strict-typed vanilla recipe
Round two of the user-style tests (six agents, real models, Cloud):
- Shutdown closed an adapter's agent but left its open tool, node and model
call — and a hand-written event.modelRequest — unclosed. Tool calls, hooks
and model calls are now tracked in the event namespace, which every one of
them is emitted through, and closed at exit in a 'leaves' phase before any
agent; a session paused on a human is skipped. Adapter agents closed at
exit also get their ProcessExit error event now.
- The unawaited-instrument() warning never fired live: the graph's nodes
arrive as roots, not under an unknown parent. A graph node arriving as a
root now warns.
- The skill's no-framework snippet failed tsc --strict on openai's union
tool types, dropped a malformed-arguments call from the trace, and lost the
ids linking tool results to their calls. Fixed in the example and the
skill, recorded as fw_tool_calls like the adapters, and the skill's block
is now type-checked against the real openai client in the vanilla suite.
- Docs: several AI SDK calls in one session() are several agents; the Next.js
warning prints at boot; ToolLoopAgent naming; LlamaIndex's default
temperature; a failed tool does not fail an error-count evaluation.
* fix(sdk/typescript): close at exit most-recent-first; name the crash that exited
Round three of the user-style tests (killed runs on LangGraph, Next.js,
Mastra, LlamaIndex, the AI SDK and a hand-built loop, checked in Cloud):
every tool, hook, model call and agent now closed on every exit path. What
was still wrong:
- Order: tools, hooks and model calls closed before any agent, so a
planner's delegate tool ended while the writer sub-agent it started was
still running. Owners now hand the exit path what they hold open, and it
closes everything in one most-recently-opened-first order.
- A crash left 'the process exited (code 1)' and no trace of the exception.
The closing messages name it (read via uncaughtExceptionMonitor, which
observes without changing how the process dies).
- A model call closed at exit carries its duration; a ProcessExit error no
longer carries the SDK's own stack as its traceback.
- Recipe and skill: tool arguments must parse to an object; the recorded
history keeps the name of each tool the model asked for. The skill covers
every exit path, the app's own records on SIGTERM, npx swallowing SIGTERM,
a missed await passing silently, Next.js request ids and cancellation, and
that a killed run's session status is still 'done'.
---------
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: chhhee10 <chetanraghuvanshi85@gmail.com>1 parent aefc062 commit f2aa5b4
209 files changed
Lines changed: 68780 additions & 36 deletions
File tree
- .github/workflows
- __tests__/ci
- docs
- reference
- sdk
- python/skill
- agents
- references
- typescript
- examples
- integration
- fixtures
- ai-4
- ai-5
- ai-6
- ai-7
- langchain-0.3
- langchain-1
- langchain-dup-core
- vendor/lc-weather-provider
- llamaindex-0.11
- llamaindex-0.12
- mastra-0
- mastra-1
- nextjs
- app
- api
- ai
- edge
- langgraph
- llamaindex
- mastra
- status
- lib
- runtimes
- types
- vanilla
- scripts
- src
- edge
- evaluator
- integrations
- test
Some content is hidden
Large Commits have some content hidden by default. Use the searchbox below for content that may be hidden.
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
484 | 484 | | |
485 | 485 | | |
486 | 486 | | |
| 487 | + | |
| 488 | + | |
| 489 | + | |
| 490 | + | |
| 491 | + | |
| 492 | + | |
| 493 | + | |
| 494 | + | |
| 495 | + | |
| 496 | + | |
| 497 | + | |
| 498 | + | |
| 499 | + | |
| 500 | + | |
| 501 | + | |
| 502 | + | |
| 503 | + | |
| 504 | + | |
| 505 | + | |
| 506 | + | |
| 507 | + | |
| 508 | + | |
| 509 | + | |
| 510 | + | |
| 511 | + | |
| 512 | + | |
| 513 | + | |
| 514 | + | |
| 515 | + | |
| 516 | + | |
| 517 | + | |
| 518 | + | |
| 519 | + | |
| 520 | + | |
| 521 | + | |
| 522 | + | |
| 523 | + | |
| 524 | + | |
| 525 | + | |
| 526 | + | |
| 527 | + | |
| 528 | + | |
| 529 | + | |
| 530 | + | |
| 531 | + | |
| 532 | + | |
| 533 | + | |
| 534 | + | |
| 535 | + | |
| 536 | + | |
| 537 | + | |
| 538 | + | |
| 539 | + | |
| 540 | + | |
| 541 | + | |
| 542 | + | |
| 543 | + | |
| 544 | + | |
| 545 | + | |
| 546 | + | |
| 547 | + | |
| 548 | + | |
| 549 | + | |
| 550 | + | |
| 551 | + | |
| 552 | + | |
| 553 | + | |
| 554 | + | |
| 555 | + | |
| 556 | + | |
| 557 | + | |
| 558 | + | |
| 559 | + | |
| 560 | + | |
| 561 | + | |
| 562 | + | |
| 563 | + | |
| 564 | + | |
| 565 | + | |
| 566 | + | |
| 567 | + | |
| 568 | + | |
| 569 | + | |
| 570 | + | |
| 571 | + | |
| 572 | + | |
| 573 | + | |
| 574 | + | |
| 575 | + | |
| 576 | + | |
| 577 | + | |
| 578 | + | |
| 579 | + | |
| 580 | + | |
| 581 | + | |
| 582 | + | |
| 583 | + | |
| 584 | + | |
| 585 | + | |
| 586 | + | |
| 587 | + | |
| 588 | + | |
| 589 | + | |
| 590 | + | |
| 591 | + | |
| 592 | + | |
| 593 | + | |
| 594 | + | |
| 595 | + | |
| 596 | + | |
| 597 | + | |
| 598 | + | |
| 599 | + | |
| 600 | + | |
| 601 | + | |
| 602 | + | |
| 603 | + | |
| 604 | + | |
| 605 | + | |
| 606 | + | |
| 607 | + | |
| 608 | + | |
| 609 | + | |
| 610 | + | |
| 611 | + | |
| 612 | + | |
| 613 | + | |
| 614 | + | |
| 615 | + | |
| 616 | + | |
| 617 | + | |
| 618 | + | |
| 619 | + | |
| 620 | + | |
| 621 | + | |
| 622 | + | |
| 623 | + | |
| 624 | + | |
| 625 | + | |
| 626 | + | |
| 627 | + | |
| 628 | + | |
| 629 | + | |
| 630 | + | |
| 631 | + | |
| 632 | + | |
| 633 | + | |
| 634 | + | |
| 635 | + | |
| 636 | + | |
| 637 | + | |
| 638 | + | |
| 639 | + | |
| 640 | + | |
| 641 | + | |
| 642 | + | |
| 643 | + | |
| 644 | + | |
| 645 | + | |
| 646 | + | |
| 647 | + | |
| 648 | + | |
| 649 | + | |
| 650 | + | |
| 651 | + | |
| 652 | + | |
| 653 | + | |
| 654 | + | |
| 655 | + | |
| 656 | + | |
| 657 | + | |
| 658 | + | |
| 659 | + | |
| 660 | + | |
| 661 | + | |
| 662 | + | |
| 663 | + | |
| 664 | + | |
| 665 | + | |
| 666 | + | |
| 667 | + | |
| 668 | + | |
| 669 | + | |
| 670 | + | |
| 671 | + | |
| 672 | + | |
| 673 | + | |
| 674 | + | |
| 675 | + | |
| 676 | + | |
| 677 | + | |
| 678 | + | |
| 679 | + | |
| 680 | + | |
| 681 | + | |
| 682 | + | |
| 683 | + | |
| 684 | + | |
| 685 | + | |
| 686 | + | |
| 687 | + | |
| 688 | + | |
| 689 | + | |
| 690 | + | |
| 691 | + | |
| 692 | + | |
| 693 | + | |
| 694 | + | |
| 695 | + | |
| 696 | + | |
| 697 | + | |
| 698 | + | |
| 699 | + | |
| 700 | + | |
| 701 | + | |
| 702 | + | |
| 703 | + | |
| 704 | + | |
| 705 | + | |
| 706 | + | |
| 707 | + | |
| 708 | + | |
| 709 | + | |
| 710 | + | |
| 711 | + | |
| 712 | + | |
| 713 | + | |
| 714 | + | |
| 715 | + | |
| 716 | + | |
| 717 | + | |
| 718 | + | |
| 719 | + | |
| 720 | + | |
| 721 | + | |
| 722 | + | |
| 723 | + | |
| 724 | + | |
| 725 | + | |
| 726 | + | |
| 727 | + | |
| 728 | + | |
| 729 | + | |
| 730 | + | |
| 731 | + | |
| 732 | + | |
| 733 | + | |
| 734 | + | |
| 735 | + | |
| 736 | + | |
| 737 | + | |
487 | 738 | | |
488 | 739 | | |
489 | 740 | | |
| |||
| Original file line number | Diff line number | Diff line change | |
|---|---|---|---|
| |||
62 | 62 | | |
63 | 63 | | |
64 | 64 | | |
65 | | - | |
| 65 | + | |
66 | 66 | | |
67 | 67 | | |
68 | 68 | | |
| |||
79 | 79 | | |
80 | 80 | | |
81 | 81 | | |
| 82 | + | |
82 | 83 | | |
83 | 84 | | |
84 | 85 | | |
| |||
0 commit comments