Reviewed: 2026-09-22 · Code-grounded.
Graph Workflows let an operator draw a directed acyclic graph of LLM calls, agent turns, tool calls, conditions and human pauses, save it, and start runs of it. The engine executes the run from the database: every node run is a row, every change is an append-only event, and the process that advances it holds no authoritative state of its own. Restarting the node loses at most the work that was in flight.
The module spans the whole stack: Client.Application/Services/GraphWorkflows/ (the parser, the state machine, the
dispatcher and the node lanes), Client.Persistence (four tables, per-column AEAD encryption), the
LocalApiRoutes.GraphWorkflows endpoint family plus GraphWorkflowRunHub, and
Client.React/src/features/graphWorkflows/ (a React Flow editor and a read-only run view).
It replaced Open Canvas. The Preview / Open Canvas visual builder was removed in the same slice that turned this feature on by default; saved canvases are converted into Graph Workflow definitions once, automatically, on the first start of the new build. That conversion is §9; read its recovery and backup guidance before upgrading a node with canvases you care about.
What it is:
- Operator-authored. A definition is a graph the operator drew. Nothing generates one, and no agent may edit one.
- Manually started. A run begins because a person pressed Start (or a caller posted to the runs route with an idempotency key). There is no schedule, no trigger, no event subscription in v1.
- Database-as-truth.
GraphWorkflowDispatcherre-reads the rows on every tick and writes back through the store. Its only in-memory state is a parsed-graph cache keyed by run id, which exists to avoid decrypting and re-parsing the pinned blob on every tick and can be dropped at any moment without changing an answer. - Durable across a restart. A run outlives the browser tab and the engine.
GraphWorkflowStartupReconcilermakes the node runs a crashed host left in flight judgeable again, exactly once, before the dispatcher's pumps start. - Acyclic.
GraphWorkflowGraph.EnsureAcyclicrefuses a cycle at save time. There is no loop construct.
What it is not:
- Not a general orchestration language. No boolean algebra in a condition (two comparisons are two edges into an
Alljoin), no expressions, no variables, no sub-workflows, no map over a collection. - Not MAF Workflows.
Microsoft.Agents.AI.Workflowsis not referenced anywhere inClient.Application; the routing here is this module's own state machine, and anAgentnode reaches the model through the same headless invocation stack a scheduled saved-agent run uses. - Not a write surface for agents. A
Toolnode may only run a built-in read-local tool that needs no approval (§4.3). A node runs unattended, so there is nobody to ask. - Not Dev Workflows. Development Workflows run a fixed template over a code work item with
gates, interventions and artifacts. Graph Workflows are free-form graphs with no work item and no artifacts, and
several of their types are deliberate copies trimmed of what has no meaning here — there is no
Blockednode-run state and noWaivededge state, because v1 has neither retry routing nor a waiving decision.
One JSON document per definition, stored encrypted as graph_json. GraphWorkflowGraph.Parse is the only entry
point and parsing is the validation: a graph that survives it is one the dispatcher can route without a second
opinion. Save time and run start call the same parser, so a graph accepted at save is a graph that will start.
schemaVersion is optional but, when present, must be 1. Anything else is refused outright.
kind says what the definition is for. A Chat graph is one the Chat page can bind a conversation to (Chat
Workflows); it may carry ChatInput nodes and publishToChat flags. An unknown
kind is refused outright, as is a chat block on a graph that is not Chat and any member of chat other than the
two above. A Chat graph without the block reads the defaults shown. The kind is denormalised onto the
definition row (graph_workflow_definitions.kind, plaintext, indexed, migration AddChatWorkflowNodes) at every
save that carries a graph, and the definition list and detail report it, so a picker filters chat workflows without
decrypting a blob. A Standard graph that names no kind is stored and returned exactly as before — the wire mapper
omits both members when they are absent.
| Member | Required | Meaning |
|---|---|---|
key |
yes | 1–64 characters of letters, digits, _ and -. Node and edge keys share one namespace. |
kind |
yes | One of Start, LlmCall, Agent, Tool, Condition, Parallel, Join, Pause, End, ChatInput, DecisionModel, by NAME. |
label |
no | Display text. Defaults to the key. |
joinPolicy |
no | All (default) or Any. A property of every node — see §2.3. |
maxAttempts |
no | Positive. Defaults to 3 for LlmCall, Agent, Tool and DecisionModel, 1 for every other kind. |
timeoutSeconds |
no | Positive. Overrides DefaultNodeTimeoutSeconds for this node only. |
position |
no | { x, y }, both numeric. Authoring metadata the runtime never reads. |
config |
no | The per-kind settings, discriminated by kind — see §4. |
config is closed per kind. GraphWorkflowGraph.ConfigMembers lists exactly what each kind reads, and a member no
node of that kind reads is an author-time error rather than a setting that silently does nothing. Writing a Tool
node's toolName on an Agent node fails the save.
Two chat members sit on several kinds and follow the same closed rule. publishToChat (Agent, LlmCall, End) marks a
node whose answer a chat-bound run posts into the conversation; it defaults to true on an End of a Chat graph
and false everywhere else, and true is refused outside a Chat graph, where no conversation could receive it.
includeAttachments (Agent, LlmCall) hands the run's chat attachments to the node (§3.8), and true is refused unless the
graph's chat.acceptsAttachments is on. Neither is read by routing; both are carried for the chat surface.
A node without a position is laid out client-side when the definition is opened
(features/graphWorkflows/models/GraphWorkflowLayout.ts). That matters for §9: imported graphs carry no positions.
| Member | Required | Meaning |
|---|---|---|
key |
yes | Same charset and namespace as a node key. Its identity — which is what makes parallel edges expressible. |
from, to |
yes | Node keys the graph declares. An endpoint the graph does not declare is a structural refusal. |
label |
no | The named outcome the source's output document reports as its branch. |
condition |
no | { path, op, value }. Absent means unconditional. |
sourceHandle |
no | Editor metadata. Read past and never stored. |
A condition is a single declarative comparison against the source node's output document, evaluated by
GraphWorkflowCondition.Evaluate. It never sees the target's input.
pathis a dot path and nothing else: property names separated by., with no wildcards, no array indexing and no functions (GraphWorkflowTokens.IsDotPath).items[0].nameis refused at save rather than saved as a property literally calleditems[0].opis one ofEq,Ne,Gt,Gte,Lt,Lte,Exists,NotExists, parsed by name (case-insensitively; a numeric token is refused, becauseEnum.TryParsewould otherwise hand back an operator no member has).valuemust be a scalar — string, number, boolean or null. An object or array has no comparison to make, and a relational operator against a boolean could never fire, so both are refused at authoring time.- Evaluation fails closed: a path the output does not carry answers
falsefor every operator exceptNotExists. An edge must never fire on data that is not there. - Numbers compare exactly —
long, thendecimal, thenBigIntegerfor integer tokens past decimal's range — so aGtover ids past 2^53 answers on the numbers rather than on a rounding artefact.doubleis the fourth and last arm (GraphWorkflowCondition.Order), reached only by the fractional and exponent tokens no exact arm reads; two that differ only past roughly 17 significant digits therefore read as equal, which is the chain's stated ceiling. A token not evendoublereads is not an ordering at all. Strings compare ordinally. A type mismatch is not an ordering and reads as "no".
A Condition node may carry a default path in its own config; its out-edges inherit it when their condition omits
one. That is authoring convenience only — the comparison still lives on the edge. An edge that resolves a path from
neither is refused, because fail-closed would otherwise make it an edge that silently never fires.
Two edges over the same (from, to) pair are legal and are how an author widens a branch. At most one of them may
be unconditional, since a second unconditional edge could only repeat the first.
This is the trap the module documents most loudly, in GraphWorkflowEnums.cs, in GraphWorkflowStateMachine.Admission
and again here: joinPolicy is a property of every node, not of Join nodes. An ordinary node with two inbound
edges joins them exactly as a Join node does.
All(the default) waits for every inbound edge to be satisfied. One dead inbound edge skips the node.Anyproceeds on one satisfied edge, but only once no sibling could still satisfy one. The parser refusesAnywith fewer than two inbound edges: one edge is anAllwritten confusingly, and none would never fire.Allover zero inbound edges is vacuously satisfied, and that is load-bearing — it is how theStartnode becomes eligible.
Pending outranks Dead under both policies. A dead inbound edge already settles what an All join will do, but
settling it while a sibling branch is still running would skip the node, and everything after it, in front of work
the run has not finished.
By the same rule, Parallel and Join are labels, not semantics. Fan-out is any node with more than one
satisfied out-edge; fan-in is the join policy. The two kinds exist because they are inline nodes that write two rows
each, which is what makes the timing of a fan-out visible in the event log at all.
Structural failures throw immediately — there is nothing useful to say about the rest of a graph nobody can walk.
Everything after them accumulates, keyed to the node or edge it belongs to, so an author fixing a canvas gets
every complaint at once (GraphWorkflowValidationResult).
Throw-first:
- exactly one
Startnode, with nothing routing into it; - at least one
Endnode, or no run could ever complete; - acyclic;
- every node reachable from
Start.
Accumulated, per element:
- a non-
Endnode with no outbound edge (a run reaching it would stop without reaching anEnd); - an
Endnode with an outbound edge; joinPolicy: "Any"with fewer than two inbound edges;- a
Conditionnode with fewer than two out-edges, or more than one unconditional out-edge; - a
Pausenode offering a decision no out-edge fires on — checked through the state machine's own routing over the document a pause actually produces, so the pre-flight rule and the routing cannot disagree; - a
ChatInputnode in a graph that is notChat, or one with no unconditional out-edge (whatever the user types, the run must go somewhere); - a second unconditional edge between one pair of nodes.
Warnings are a second list, and they never refuse. GraphWorkflowGraph.Warnings is computed on the first ask
rather than during the parse — only the validate endpoint asks, and every dispatcher tick parses — and nothing in it
reaches GraphWorkflowValidationException. A graph that warns saves, validates as valid, and runs. There are
two warning kinds in v1, and Warnings is their concatenation.
The first: a node whose inbound edges all leave a Pause receives the decision document rather than the
content that was approved (§4.6). It is keyed on that node rather than on the pause, because that is the node which
loses the content and so the node an editor draws the badge on, and one warning is raised however many pauses reach
it. The sentence names the pause's nearest non-Pause ancestor only when that ancestor is unique and is not a
Condition: two candidates means the pause is fed by mutually exclusive branches, and edges from both would leave an
All successor waiting on the branch that was never taken, while a Condition cannot be named because the edge it
would ask for is that node's second unconditional out-edge, which the parser refuses — advice that turns a warning
into a hang or an error is worse than the generic sentence.
A successor whose joinPolicy is Any is exempt, whatever its kind, and that is a correctness rule rather than
a taste one. The advised edge is unconditional, so it stays satisfied when every approval is rejected — an Any node
would then be admitted on the content edge alone and run the branch the rejections were meant to stop. Collecting the
decision documents is what such a node is for, so there is nothing to warn about. An All successor waits for the
approval edges too, so it keeps both the warning and the advice.
The second, on an Agent node whose responseJsonSchema asks for something the runtime it will run on does not
enforce (GraphWorkflowGraph.ResponseSchemaWarnings). It fires on three things: a keyword the
Microsoft.Extensions.AI.OpenAI strict-schema transform relocates into description, a declared property missing
from required, and an object that declares properties and omits additionalProperties — an object that sets
that key explicitly, true included, is silent. The grammar is then built from the rewritten schema, so
maxLength: 3 survives only as a hint the model may read and nothing enforces, while a property the author left
optional comes back mandatory. Structure — type, enum, required, the object shape — survives the transform,
which is why none of it is warned about. One warning per node carries whichever of the three apply, and each
list names at most three before counting the rest, because this is a sentence and not an inventory:
Node 'agent' declares a response schema the runtime rewrites before it becomes a grammar: it drops 'maxLength'
rather than enforcing it, requires every declared property ('summary', 'notes' are optional here) and forbids
additional properties.
Since S7 this warning is narrowed at the service, not at the parser. The rewrite it describes is the OpenAI
adapter's, and the llama.cpp lane no longer suffers it (§4.2): a node whose model llama-server serves receives the
schema as authored. The parser cannot tell which nodes those are — it has no route to the model-to-provider map — so
it keeps raising the warning for every Agent node and GraphWorkflowDefinitionService.ValidateAsync drops the ones
whose pinned model resolves through ILocalModelProviderResolver to llamacpp. The filter is per warning
rather than per node: a node that also earns the pause-context warning above keeps that one. A node with no model
pin keeps its schema warning too — what it inherits is decided at run start, and a warning nobody needed is the
cheaper error. Nothing else moved: the parser's rule set, the warning sentence and the DTO are all unchanged, and a
warning still never blocks.
It warns rather than refuses because a schema is still useful with the constraints in it, the transform is the adapter's business and could change, and every one of these graphs runs. The point is that the author stops believing the parts that do not hold.
The schema is walked breadth-first over exactly the members the transform itself descends — properties,
additionalProperties, items, anyOf, oneOf, allOf — so a constraint under items is found and one parked in
a $defs or definitions pool is correctly ignored, since the transform never reaches it either. A Start node's
inputSchema never warns: nothing compiles it into a grammar, on any runtime. Like every warning it is non-blocking, and the
editor's validation strip renders it through the same generic channel as the first kind, so neither the DTO nor the
SPA needed a new shape for it. What it is warning about is §4.2.
The node cap runs first of all, ahead of every rule above. MaxNodesPerDefinition reaches the parser as an
argument rather than a dependency (GraphWorkflowGraph.Parse(graphJson, maxNodes), defaulted to no cap so the parser
stays testable without a container; GraphWorkflowGraphContract.ValidateAndCountNodes passes the option through), and
it is checked against the declared length of the nodes array before a single node is read or an edge walked. A
cap applied after the parse would bound nothing about the parse that produced it: only the 1 MiB body limit stood
between a request and a chain of thousands of minimal nodes, and the acyclicity walk used to spend one stack frame per
node, so a deep enough chain overflowed the thread-pool thread's stack — a process kill, not a 400. That walk is now
iterative over an explicit stack, which is also what keeps the deliberately uncapped re-parses of a stored graph
safe (§3.4). The refusal for a cycle names every node on it in walk order (a -> b -> c -> a), starting and ending at
the node the walk came back to.
The eight-node graph below is the module's canonical shape — Start → Agent → Condition → { Pause | Tool } → Parallel → Join → End. It is the React tests' eightNodeGraph fixture
(features/graphWorkflows/test/GraphWorkflowFixtures.ts) with every node's position and analyze's
timeoutSeconds: null left out for reading — the graph is otherwise the same one.
{
"schemaVersion": 1,
"nodes": [
{ "key": "start", "kind": "Start", "label": "Start",
"config": { "inputSchema": null, "defaultInput": null } },
{ "key": "analyze", "kind": "Agent", "label": "Analyze", "maxAttempts": 3,
"config": { "agentDefinitionId": null,
"instructions": "Summarise the request and say whether it needs a human review.",
"model": null, "reasoningEffort": null,
"responseJsonSchema": { "type": "object",
"properties": { "requiresReview": { "type": "boolean" } } },
"includeUpstreamOutputs": true } },
{ "key": "check", "kind": "Condition", "label": "Needs review?",
"config": { "path": "output.json.requiresReview" } },
{ "key": "review", "kind": "Pause", "label": "Human review",
"config": { "prompt": "Approve the analysis?",
"allowedDecisions": ["Approve", "Reject"], "requireComment": false } },
{ "key": "lookup", "kind": "Tool", "label": "Read file", "maxAttempts": 3,
"config": { "toolName": "read_file", "arguments": { "path": "notes.md" },
"argumentBindings": { "path": "output.json.path" } } },
{ "key": "fanout", "kind": "Parallel", "label": "Both", "joinPolicy": "Any", "config": {} },
{ "key": "merge", "kind": "Join", "label": "Merge", "joinPolicy": "All", "config": {} },
{ "key": "done", "kind": "End", "label": "Done", "joinPolicy": "Any",
"config": { "outcome": "completed", "resultPath": null } }
],
"edges": [
{ "key": "e1", "from": "start", "to": "analyze" },
{ "key": "e2", "from": "analyze", "to": "check" },
{ "key": "e3", "from": "check", "to": "review", "label": "yes",
"sourceHandle": "true", "condition": { "op": "Eq", "value": true } },
{ "key": "e4", "from": "check", "to": "lookup", "label": "no",
"sourceHandle": "false", "condition": { "op": "Ne", "value": true } },
{ "key": "e5", "from": "review", "to": "fanout", "label": "approved",
"sourceHandle": "Approve",
"condition": { "path": "output.decision", "op": "Eq", "value": "Approve" } },
{ "key": "e6", "from": "lookup", "to": "fanout" },
{ "key": "e7", "from": "fanout", "to": "merge" },
{ "key": "e8", "from": "merge", "to": "done" },
{ "key": "e9", "from": "review", "to": "done", "label": "rejected",
"sourceHandle": "Reject",
"condition": { "path": "output.decision", "op": "Eq", "value": "Reject" } }
]
}Two details in it are the rules of §2.4 doing their job, and both are easy to get wrong:
e9exists becausereviewoffersReject. APausenode whose allowed decision has no out-edge that fires on it is refused at save — answering it would strand the run.fanoutanddonedeclarejoinPolicy: "Any". Thechecknode's two branches are mutually exclusive, so under the defaultAlljoin each would wait forever for a branch that was never taken, and the run could never reach an end.
e3 and e4 carry no path: they inherit check's config.path.
POST graph-workflows/definitions/{definitionId}/runs answers 202 with the run id.
IGraphWorkflowRunService.StartAsync validates, commits a durable intent and signals the dispatcher, which advances
the run out of band — so the run legitimately reads Pending when the answer lands.
The caller's requestId is the idempotency key. The lookup before the insert is a fast path two concurrent callers
can both pass; the unique index on request_id is the real guarantee, and a reused id naming a different
definition is refused rather than answered with a run the caller never asked for.
At start the run pins its own copy of the graph (graph_json and graph_hash on the run row), and
StartRunAsync commits the run row, one Pending node run per graph node and the run.created event in one
transaction, re-reading the definition inside it so a delete racing the start cannot leave an orphan run. Every node
of the pinned graph therefore has a row from the moment the run begins — a Pending row means "not yet judged", not
"not yet reached". Editing or deleting the definition afterwards never rewrites a run.
Two checks happen again at start rather than being trusted from save time, because the world moves in between: the
graph is re-parsed with the same parser, and every Tool node's tool is re-checked against the live catalog (§4.3).
The run also refuses a graph over MaxNodeRunsPerRun and an input over MaxRunInputBytes.
Pending → Running → WaitingForApproval → Completed | Failed | Cancelled, plus Cancelling
(GraphWorkflowRunStatus). There is no Interrupted: runs auto-resume after a host restart and only node runs
reconcile. Cancelling exists because cancel is fire-and-forget — the endpoint commits an intent and returns 202,
and live node runs drain first.
The run's status is recomputed from scratch at the end of every tick, never accumulated
(GraphWorkflowStateMachine.Recompute). It is denormalized so a reader can answer "what is this run doing" without a
join. The recomputation is graph-aware, and the ordering matters:
- A failed node run outranks everything: the run is
Failed, carrying that node's own failure class and reason. When several failed, the one with the lowest node key ordinally is picked, so two readers of the same run cannot disagree. - Otherwise, if a terminal node of the graph (one no edge leaves)
Succeeded, the run isCompleted. - Otherwise the run is
Cancelled, with a reason naming the ends it did not reach and what became of them. A run whose tail was skipped, or whose rejection routed down a branch that skipped the remainder, therefore readsCancelledrather thanCompletedlike a run that did its job.
While node runs are still live, WaitingForApproval outranks Pending: every node run exists from the start, so
there are almost always Pending rows, and reading those as Running would report a run blocked on an unanswered
pause as busy — the one thing the two statuses exist to tell apart.
GraphWorkflowFailureClass records why: None — the default, carried by everything that has not failed — then
NodeFailed, Timeout, AttemptsExhausted, OutputTooLarge, GateRejected, ValidationFailed, Cancelled,
Interrupted, CapacityRejected. GateRejected narrows where a reader should look; it
is not a causal proof, because a rejection can route into a branch that then runs perfectly well.
Pending → Queued → Running → Succeeded | Failed | Skipped | Cancelled, plus WaitingForApproval
(GraphWorkflowNodeRunStatus). Queued and Running are separate because an admitted node run is not yet
executing. There is no Blocked: v1 has no retry routing, so there is no retries-exhausted intervention state to
park a row in.
GraphWorkflowStateMachine.Admission answers Wait, Eligible or Skip for a Pending row from its inbound edge
states alone. An inbound edge is Satisfied when its source succeeded and its condition fired, Dead when the
source settled any other way or its condition did not fire, and Pending until the source is terminal — a source
with no row yet is a wait, not a refusal.
A skipped row records one cause, and prefers a branch that broke or was skipped over one a condition merely
routed past: a Condition node taking its other branch is the graph working, not news
(GraphWorkflowStateMachine.SkipReason).
Retry is in place: Failed → Pending is the one edge out of a terminal node-run status. A failed row under both
the node's maxAttempts and the run's MaxTotalAttempts goes back to Pending with the attempt incremented, in one
atomic write, and the node.retried event carries the failure the row cleared. Only NodeFailed, Timeout and
Interrupted are retryable (GraphWorkflowFailures.IsRetryable) — an over-cap document, a refused gate, a cancelled
run, a graph that no longer declares a node and a capacity refusal (CapacityRejected: only the operator ejecting a
model or picking a loaded one changes it) all produce the byte-identical answer next time.
A consequence worth stating plainly: a node declaring maxAttempts: 1 reports AttemptsExhausted on its only
attempt, because the node's budget is genuinely why nothing will try again. What actually went wrong survives on the
row's reason and on its node.failed event.
GraphWorkflowDispatcher is one loop with two pumps — a signal channel (bounded, drop-on-full, because a signal is a
latency hint) and a sweep every DispatchIntervalMilliseconds over every live run. A dropped signal costs at most one
interval of latency, never correctness. The first sweep runs immediately at startup rather than after an interval, so
recovery does not pay for one.
Every dispatch-side status write happens inside a serialized AdvanceOnceAsync call. A lane's work produces a
pollable result and never transitions a row itself; the only other writers to a run are the human command paths.
One tick, in order, and the order is load-bearing:
- Steer (§3.7). Judge every unjudged steer entry: a row still
Queued/Runningon the steered attempt and invocation goes back toPendingon the same attempt; any other entry is recorded ignored. Before the poll, so the poll's superseded-entry sweep cancels the old turn in the same tick. - Poll. Ask every lane what became of the work it was driving and settle what landed — before anything reads the
rows for a decision, or the run judges its graph against a row that is only still
Runningbecause nothing asked. A row its lane had nothing to say about is offered to its deadline instead. - Drain (only while
Cancelling). A drain admits nothing. Every terminal is reached through it or through the "nothing is live any more" recomputation, because writing a terminal over live node runs would strand them under a run no tick looks at again. - Retry the failed rows that still have budget.
- Admit the eligible
Pendingrows, and skip the ones every path into which is dead. - Publish (chat-bound runs only, §3.6): every
Succeededrow whose node haspublishToChatand nopublished_message_idyet becomes one chat message. After admit, which settles the inline kinds (End among them), and before the recompute that may end the run. - Recompute the run's own status, against the version it was read at.
The concurrency cap counts working runs, not parked ones. Admission of a Pending run asks
CountActiveRunsAsync whether MaxConcurrentRuns are already executing. A run whose live rows wait on a person —
at least one WaitingForApproval row, none Queued or Running — is parked and holds no slot, so four chats
waiting on their users cannot freeze every new start at Pending behind a 202. Cancelling always counts. A parked
run that resumes does not re-pass the admission gate (accepted: the lanes still bound real work). This applies to a
Standard run parked on a Pause too.
Deadlines are re-derived from the row every tick — started_at_utc plus the node's timeoutSeconds or
DefaultNodeTimeoutSeconds — never armed in memory, so they survive the restart that would otherwise leave a node run
bounded by nothing. GraphWorkflowDeadline.Grace adds 30 s before the run ends a row itself: the lane bounds its own
turn by the same number from a slightly later moment, and ending the row the instant the number is reached would race
a better answer and sometimes win by milliseconds.
GraphWorkflowStartupReconciler is an IHostedService registered before the dispatcher, so its pumps cannot
admit a row recovery has not judged. It reads the interrupted set — exactly Queued ∪ Running, which the store
scopes and the reconciler never widens. WaitingForApproval is deliberately outside it: it is a durable human wait
rather than in-flight work, and a reconciler that took it would destroy every pause and chat input on the node on every boot. The
reconciler hands its verdicts to the store
to apply in one transaction. Exactly-once survives a crash during recovery: a host that dies before that commit
leaves the rows as it found them, and the next boot judges them from the same evidence. Recovery makes at most three
passes, and the last one settles whatever is left rather than walking away from it.
The verdicts:
- A
Queuedrow, or aRunninginline row, collapses toPendingwithout touchingAttempt. Neither is a failure, so neither costs an attempt. - A
RunningLlmCall,Agent,DecisionModelorToolrow is failedInterrupted. The work was an in-process task with no durable handle, so its partial output died with the host. Recovery never re-attempts it; the dispatcher's retry stage does on its first tick, if and only if the node and run budgets allow. The class written is the plainInterruptedrather than anythingGraphWorkflowFailures.Classifywould decide, because recovery deliberately never parses the run's graph and so cannot see the node's attempt cap — the retry stage, which does, classifies on that first tick.
That is the same "never resume a provider stream, start a fresh attempt instead" posture Development Mode records in ADR 0001, applied without that design's replacement-attempt rows: a graph workflow node run retries in place with the attempt incremented, because per-attempt history lives in the event log rather than in a second row.
The reconciler touches no run row. The dispatcher's first sweep — PumpSweepAsync runs one immediately rather than
waiting out an interval — recomputes the run's status from the rows recovery left behind.
A run started from the Chat page (POST graph-workflows/conversations/{conversationId}/messages, Chat
"Chat workflow mode") is bound to that conversation: graph_workflow_runs.conversation_id is a real foreign key to
conversations with ON DELETE SET NULL, and trigger_message_id names the user message that started it (migration
AddChatWorkflowBinding). A partial unique index on conversation_id over the four live statuses
(ux_graph_workflow_runs_live_conversation) makes one live run per conversation a database rule; the store answers the
losing insert with GraphWorkflowRunBusyException. IGraphWorkflowRunService.StartAsync has an overload taking a
GraphWorkflowRunBinding; an unbound run (the Graph Workflows page) publishes nothing. While the newest bound run is
live (parked included), a normal chat send into the conversation is refused server-side with conflict type
GraphWorkflowRunLiveInConversation, not only by the Chat page's composer lock (Chat "Chat workflow mode"). Every Agent of a Chat graph sees
the chat message in its prompt (§4.2, "Conversation request"). A conversation purge
(ConversationFootprintPurge) unbinds the conversation's runs explicitly and never deletes one.
Publishing is an outbox pass in the tick (§3.4 step 5), GraphWorkflowChatPublisher: one query for the candidates
(ListUnpublishedNodeRunsAsync: Succeeded and published_message_id IS NULL), the node's publishToChat read off
the pinned graph, an insert-if-absent of the assistant message under the deterministic id
GraphWorkflowChatIds.PublishedMessage(runId, nodeKey, attempt), then MarkNodeRunPublishedAsync — a compare-and-set
that stamps graph_workflow_node_runs.published_message_id and appends node.published (detail { messageId }) in one
transaction. The stamp does not bump the run's version — it moves no run state, so no run-level write should lose a
race to it. Either crash order ends as one message: an insert that landed before a crash is found by its id, and a
stamp that landed is never repeated. The message is written with a thin run envelope in the same transaction, so the
restart reconcile never backfills it as a chat run. A publish failure is logged and never fails the node; while one
fails, the tick skips the recompute (which could be the one that ends the run, after which nothing ticks it again)
and counts as work so it re-signals, at most three ticks per run (GraphWorkflowDispatcher.MaxPublishRetries, in
memory) — after that the run completes without the message. A run that is cancelled drains without publishing.
Operator ruling D4 (2026-09-23): the running Agent or LLM Call node of a chat-bound run can be steered. POST graph-workflows/runs/{runId}/nodes/{nodeKey}/steer with { operationId, message } goes through
IGraphWorkflowChatService.SteerAsync, which writes the text as a user message of the run's conversation under the
deterministic id GraphWorkflowChatIds.SteerMessage(operationId, runId, nodeKey) and then calls
IGraphWorkflowRunService.SteerAsync. An unknown node key is a 404 before anything is written. When the run service
refuses, not-found included, the message is removed again. The service commits intent only. It appends
{ operationId, message, atUtc, attempt, invocationId, steeredBySubject, applied: null } to the encrypted
graph_workflow_node_runs.steering_json (migration AddGraphWorkflowSteering, AAD purpose
graph_workflow_node_run_steering_json), signals, and answers 202 with the run detail. The append is a
compare-and-set that re-checks the run, the row and the cap inside its transaction. It does not bump the run version.
Refusals: 400 for a node that is not Agent or LlmCall, for a run with no conversation, and for a blank or oversized
message (the chat-input answer cap, min(MaxMessageSizeKb × 1024, MaxOutputJsonBytes / 2)). 409
GraphWorkflowSteerLimitReached for a row already holding MaxSteersPerNode entries. 409 GraphWorkflowRunConflict
for a Pending, Cancelling or terminal run, for a row that is not Queued/Running, and for a reused
operationId with a different message or person.
The same operationId with the same message is a replay and answers 202 again. Idempotency is per entry, not per
attempt, because a steer never changes the attempt.
The dispatcher applies it (§3.4 step 0), because the store's node move has no from-status predicate and only
inside the advance gate is "still running" still true when the write lands. A row still Queued/Running on the
entry's attempt, and on its invocationId when the steer saw one, moves to Pending with node.steered, and its entries are stamped applied: true in the same transaction. Every
applied entry gets its own node.steered event (two steers between ticks are one reset and two events). Any other
unjudged entry is stamped applied: false with node.steer-ignored. Both details are { operationId, message, attempt }, so the activity feed renders a running node's steer from the event alone. A Cancelling run
ignores them all, and so does a terminal one: a steer can commit after the apply pass of the tick that ends the run,
so the terminal early return in AdvanceCoreAsync judges it on the tick the steer's own signal brings. The poll's ForgetSupersededAsync then discards the old flight, and the invocation executor's
discard hook cancels its invocation. The next admit re-runs the node on the same attempt: no attempt is spent,
MaxTotalAttempts is not charged, and the restart reconciler is unchanged, since a Pending row is not interrupted
work. The re-run's seed prompt ends with an ## Operator steering section listing the applied entries, oldest first
(§4.2, §4.2a). Each reset restarts the node deadline, so MaxSteersPerNode is the bound. Steering text is not
counted against MaxRunInputBytes: each entry is bounded by its own per-entry byte cap (above) and a row by
MaxSteersPerNode. Cross-node routing stays
absent (register 22, D6(a)).
A chat send to a graph whose chat.acceptsAttachments is on carries references only: run.input.attachments is
[{ fileId, name, kind: "text"|"image", bytes }] (bytes is the upload's size), checked with the rest of the run input
against MaxRunInputBytes before the user message is persisted. Any other graph answers 409
GraphWorkflowAttachmentsNotAccepted (Chat "Chat workflow mode").
Content is read at attempt time, and only by an Agent or LLM Call node with includeAttachments: true on a
chat-bound run (run.input.conversationId). GraphWorkflowInvocationExecutor resolves each reference through
IConversationUploadedFileStore:
- Text — the cached extracted Markdown, composed by the chat's own
ConversationAttachmentContextComposerinto a## Attachmentssection after the prompt (and before any steering section). The section opens with the chat's notice that the content is untrusted DATA, not instructions. Each file is wrapped byUntrustedContentFraming.WrapDocument, with its name inside the fence asfile:metadata. The marker's nonce is keyed on the server-secret per-conversation seed (IUntrustedContentFenceSeedProvider) plus the run id and node key, so a document cannot forge the closing marker. Bodies are budgeted before wrapping atMaxRunInputBytes, counted in characters (the composer's unit). Truncation therefore never cuts a closing marker, and the composer's truncation notice follows. A file whose status is notExtracted, or whose Markdown is empty, is skipped, below. The Markdown is read only for anExtractedrow, because a.mdupload's raw bytes share the Markdown blob's path. - Image — attached as a
ConversationImageParton the seed turn throughIChatTurnContextBuilder.BuildImageContextAsync, so the chat'sMaxImageAttachments/MaxImageAttachmentBytescaps apply (over-cap images are dropped with a warning). Only when the node's effective model resolvesSupportsVision; otherwise the node failsValidationFailedwith a reason naming the node, the file and the model. An LLM Call carries images the same way. Animagereference whose file's extraction status is notImageis skipped before the vision check, so it never reachesBuildImageContextAsync, which would drop it silently. - Skipped — a file no longer in the conversation when the attempt starts, and the unusable text or image files above, are skipped. The same attempt-time resolution
that reads the content decides this, so the notice can never disagree with the prompt. The node's output document
gains
attachmentsSkipped: [name…](omitted when nothing was skipped), the executor logs the count at Information, and the turn still runs.
A node without the flag, a Standard graph and a run without attachments send a prompt byte-identical to one before
attachments existed; every Agent of a Chat graph still sees the Attachments: <names> line of its conversation request (§4.2).
Every kind's output goes through one writer, GraphWorkflowDocuments, which composes the common envelope, derives
the branch and enforces the size cap. No executor composes a document itself, because a second implementation of
any of those is a way for an executor to disagree with the routing the dispatcher will do a moment later.
The output envelope — what an edge condition reads:
{ "status": "succeeded" | "failed" | "skipped",
"attempt": 1,
"branch": "approved" | null, // the label of the first CONDITIONAL out-edge that fires; null when none did
"output": { /* per-kind, below */ } }The input document — what an executor is handed:
{ "run": { "input": /* the run's start payload */ },
"upstream": { "<nodeKey>": /* that node's output envelope */ },
"input": /* the single satisfied predecessor's whole envelope, the upstream map when there
is more than one, or null when there is none — which is what Start sees */ }branch is written even when null, so a reader can tell "no branch fired" from "this document predates branches".
An unconditional out-edge never names a branch: it accepts everything, so it says nothing about which way the run
went. The composed document is capped at MaxOutputJsonBytes in UTF-8 bytes — a character count would let astral
text through at four times the cap — and a node whose document exceeds it fails OutputTooLarge, which is not
retryable.
Pause and ChatInput park on a person (GraphWorkflowPauseExecutor); Agent, LlmCall and DecisionModel share
the model-invocation lane (GraphWorkflowInvocationExecutor); Tool has its own lane.
Start, End, Condition, Parallel and Join are inline (GraphWorkflowInlineExecutor): their work is a
pure function of rows the tick has already read, so they run inside the tick with no Queued hop. They still write
two rows each, which is what makes the timing of a fan-out visible in the event log.
Config: inputSchema, defaultInput — both optional raw JSON, and both authoring metadata in v1 (the run's input is
validated for size, not against the schema).
Output: { "input": <the run's start payload> }, handed to everything downstream.
One headless saved-agent turn, on the model-invocation lane shared with LLM Call and sized at MaxConcurrentRuns
(GraphWorkflowInvocationExecutor). It drives IInvocationRunner from the tick and never inside it; the turn is an
in-process task with no durable handle, which is exactly why an interrupted Running row is failed rather than
resumed (§3.5). InvocationId is written on the row as the correlation id in the node logs for a turn nothing else
survives.
The turn's contents come from RunSavedAgentHandler, the node's other unattended caller of this stack, and five of
its rules carry over: the locality gate runs before capacity, the capacity reservation is disposed on every
terminal path, approval-required tools are stripped from the offer (and the web tools by name, §4.3), IsUnattended is set, and the terminal state is
read off a StrongBox<T> the state-changed handler fills rather than off a return value the runner does not have.
The executor is a singleton — the lane and its slot count are the node's and outlive both a tick and a DI scope —
so every scoped collaborator is resolved inside the task body from its own scope: the scope the tick handed the store
is long gone by the time the turn lands.
| Config member | Meaning |
|---|---|
agentDefinitionId |
Optional saved agent to bind. Null runs an unbound persona from instructions alone. |
instructions |
Required. The seed user turn. |
model |
Optional model pin. Travels as written and is matched against the catalog at run start, exactly as an agent definition's own pin is, so a graph does not become unsaveable because a model was uninstalled after it was authored. |
reasoningEffort |
Optional, one of none, low, medium, high. Checked at save — its vocabulary is closed and cannot go stale. |
responseJsonSchema |
Optional. Must be an object schema ("type": "object"), because the parsed answer lands at output.json and a condition reads a property off it. |
includeUpstreamOutputs |
Defaults true. Inlines the upstream map into the prompt, budgeted at MaxRunInputBytes and truncated with an explicit marker rather than silently. |
In a Chat graph every Agent's seed prompt also carries the chat request, whatever includeUpstreamOutputs says:
after the instructions and before the upstream map, a ## Conversation request section with run.input.message (same
MaxRunInputBytes budget and truncation marker) and, when the run input lists attachments, one Attachments: a.pdf, b.png line (names only; the content reaches an includeAttachments node, §3.8). Agent nodes have no inputBindings, and after a router the upstream is
the router's choice rather than the message, so without it an agent downstream of a DecisionModel or Condition never sees
what the user asked (live-round finding, 2026-09-23). A Standard graph's prompt is unchanged byte for byte.
A steered row's prompt ends with an ## Operator steering section (§3.7): the row's applied steers, numbered,
oldest first, after everything above, the ## Attachments section (§3.8) included. An LLM Call gets the same section after its bound prompt. A row nobody steered
sends a prompt byte-identical to one before steering existed.
Output: { "text": …, "json": … | null, "usage": { inputTokens, outputTokens, totalTokens, reasoningTokens, durationMs, finishReason, model } }. Every usage member is nullable: the runner reports what its provider gave it,
and no provider reports all of them.
What a response schema actually enforces. llama-server compiles it into a real GBNF grammar and applies that
grammar from the first output token, so the structure is guaranteed: the object shape, the declared types, an
enum's member list, the required keys. The compiled root rule leaves the model's <think> block optional and
unconstrained and forces the schema only on what follows, so a reasoning model is not fighting the grammar.
Since S7 the value bounds are guaranteed too, on llama-server. minLength, maxLength, pattern, format, the
numeric minimum and maximum, the item counts and the rest now reach llama.cpp as written, and so does the author's own
required list — a property left optional stays optional, and an object that omits additionalProperties stays open.
DeferredLlamaServerChatClient.ApplyResponseSchemaPassthrough writes the authored schema onto the request body
itself, and the MEAI OpenAI adapter fills its own response format in only with ??=, so its unconditional
strict-schema transform is never invoked on this lane. The one exception is a repetition bound above
LlamaGrammarToolSchemaCompatibility.MaxGrammarRepetitionBound (1024), which is still stripped because llama.cpp's
GBNF converter cannot compile it into a grammar at all — see §8.
On every other runtime the old rule stands. The strict-schema transform relocates twenty-two value keywords into
the schema's description string on the way out of the .NET process, where they are advice to the model rather than a
constraint (a maxLength: 3 produced a 1302-character field in the S6 live round), marks every declared property
required, and closes an object that has properties and omits additionalProperties. There, validate a value bound
downstream — in an edge condition or the consuming node — and never in the schema alone. You do not have to notice
this unaided: saving or validating a graph raises the non-blocking warning of §2.4 on an Agent node whose schema asks
for it, and that warning is now raised only for nodes that are not pinned to a llama-server model. See
docs/agent-knowledge.md §2 for the evidence and the --verbose recipe that shows the compiled grammar.
A node declaring a response schema must come back a JSON object. A parse failure fails NodeFailed — the
retryable class, since a re-ask under the same grammar can land where one attempt did not — naming the finish reason,
because a truncated answer is still Completed and length is the common cause. There is deliberately no salvage
path: grammar-constrained output carries no fences, and stripping some would quietly mask a broken grammar. See §8
for what a response schema costs at the llama.cpp grammar layer.
An Agent node is node-local only. If the node's effective model resolves to a cloud model the node run is
refused ValidationFailed before any capacity is reserved: an unattended run must never hand the operator's
content to a cloud provider. Approval-required tools are likewise stripped from the offer, and the turn runs
with IsUnattended set — the same posture a scheduled saved-agent run holds.
The runner's own watchdog reports a timeout as a failed terminal, which is why the executor classes it Timeout
rather than NodeFailed — a node that ran out of time deserves a different answer from one whose provider said no.
Provider errors never reach the row: the reason says "see the node logs" and the detail stays there.
One logical chat-model invocation per node attempt, through GraphWorkflowInvocationExecutor.RunLlmTurnAsync and the shared
ILocalChatRuntimePackageBuilder / IInvocationRunner stack. The node does not resolve an agent definition or add
a persona, tools, skills, memory, or orchestration. Existing bounded transport retries before the first output
remain enabled; graph retries are separate attempts governed by maxAttempts. Tool-free requests bypass the
function-invocation loop, so unsolicited tool-call content cannot trigger a recovery round.
The package's explicit OmitSystemPrompt flag carries absence through validation, configuration hashing and
InvocationAgentFactory.BuildSeedMessages. It defaults to false, so existing chat, Agent and benchmark callers
retain their non-empty system-prompt requirement and their existing configuration hashes.
The LLM Call package also requires node-managed llama routing for the selected model throughout dispatch. The shared runtime bypasses cloud selection for that request and rejects a provider remap before choosing a local client, including a cached client. This keeps a settings change between validation and inference from sending the call externally, without locking provider settings for the duration of generation.
| Config member | Meaning |
|---|---|
model |
Optional installed node-managed GGUF chat model. When omitted, the local-default resolver chooses one. Cloud, external, Ollama, uninstalled and non-chat models are refused before inference. The chosen model is not automatically swapped. |
systemPrompt |
Optional authored system instructions. Empty means no system-role message; no default assistant prompt is substituted. |
prompt |
Required user instructions for this call. |
inputBindings |
Optional object mapping names to dot paths in the node's input document. Values form a named JSON data block alongside the prompt; there is no placeholder expansion or implicit upstream context. |
reasoningEffort |
Optional, using the same vocabulary as the Agent node. Actual support depends on the selected model. |
responseJsonSchema |
Optional object schema, with the same response handling and provider limitations as Agent output. |
samplingOptions |
Optional inference overrides. Unset values use runtime defaults. The editor keeps these in a collapsed Advanced section. |
Bindings use GraphWorkflowDocuments.Resolve. A missing path fails the node before inference; an explicit JSON
null is a valid bound value. run.input.question reads the run input, and
upstream.retrieve.output.result reads an incoming Tool result. input.output.result is a shortcut only when
exactly one incoming predecessor is satisfied. The upstream map contains satisfied immediate predecessors,
not every ancestor: add a direct dependency edge to consume an earlier result, taking the consumer's join policy
into account. Bound data is size-limited and is not silently truncated into a different JSON value.
Sampling overrides include temperature, topP, topK, minP, maxOutputTokens, seed, repeatPenalty,
repeatLastN, presencePenalty, frequencyPenalty, stop, and numCtx. The last is a request/prompt budget
ceiling; it cannot resize the llama-server context window fixed when the model was loaded.
Output uses the Agent-compatible { text, json, usage } payload inside the normal node envelope. A requested
structured response must parse as a JSON object; otherwise the attempt fails. This is not a separate full JSON
Schema validation pass. A following Condition can route on output.json.needsReview without parsing prose.
For example, Start → Tool → LLM Call → End can bind document to input.output.result and use the prompt
"Summarize the bound document and retain the technical findings." Only the named document is supplied to the
model. A second LLM Call can consume the first one's input.output.text the same way.
One built-in tool call, on its own lane, through IToolInvocationService (GraphWorkflowToolExecutor).
| Config member | Meaning |
|---|---|
toolName |
Required. |
arguments |
Optional object of literal arguments. |
argumentBindings |
Optional map of argument name → dot path, resolved against the node's input document. A binding wins over a literal of the same name. |
allowedUrls |
web_fetch only, and required there: the link allow-list (below). Any other tool carrying it is a save error. |
Output: { "result": … } — embedded as JSON when the tool answered with an object or an array, and as a string
otherwise. That one try-parse is what lets a downstream Condition, which passes its predecessor's output through
verbatim, dot-path into a structured answer. Stated rather than hidden: plain text that happens to be a JSON object
is embedded as JSON, and the tools that do that mean it.
A Tool node passes two gates, not one (GraphWorkflowToolGate, ruling D6 and
ADR 0006):
- the tool must be a built-in tool in the
ToolCategory.ReadLocalenvelope, and - the composed approval policy must answer that it needs no approval.
Both are asked of the same catalog at definition save and again at run start, and the run-start answer wins. A tool that was read-only when the graph was saved does not execute after a policy tightened. A tool outside the envelope is an error, never a warning: a workflow node runs unattended, so a write, execute or approval-gated tool has nobody to ask. Errors are keyed by node key, so the editor draws them on the offending node.
The one web exception (ADR 0017 decision 6; ADR 0006 is its
basis and is not amended). While web access is on, a Tool node may name web_fetch if it carries a non-empty
allowedUrls list of absolute http(s) URL prefixes (same scheme, host and port; a path prefix on a segment boundary).
The list, written before the run, is the consent an unattended node cannot ask for, so there is no pause and no result
review. The gate checks the list's shape and any literal url argument at save and start; a bound url,
usually upstream model output, is checked at dispatch, which is the real gate: GraphWorkflowToolExecutor calls the
fetch service itself with the node's list, and the first URL and every redirect hop must match. Private and local
addresses stay blocked even when listed. IToolInvocationService still refuses web_fetch (it is Network), so no
other caller reaches it this way. A URL outside the list, a blocked address and web access switched off fail the node
ValidationFailed; any other refusal about the page (HTTP error, unsupported type, the fetch's own time budget)
succeeds carrying the tool's { "error", "message" } answer, as read_file's own refusal does. web_search is never
allowed on a Tool node, and no Agent node — bound agent or default persona — is offered either web tool: both are
stripped from the offer by name.
The lane itself is a GraphWorkflowInFlightLane, the same registry the Agent lane uses, and the two differ only in
what they run: dispatch to a queue, settle on the poll, a stop that answers no on a repeat, forget what a retry
superseded. A tool call passes one queue where an agent turn passes two — there is no node-wide invocation lease
to wait for after the lane slot, so the row reads Running the moment the lane hands back an entry. The lane is
therefore what bounds the fan-out: a tool call has no global bottleneck of its own, and a Parallel node feeding two
hundred search_knowledge_base nodes would otherwise fire all of them at once. It is sized on MaxConcurrentRuns
rather than on a knob of its own — the same "how much of this node may be busy at once" question — and is worth a
second option only once the two need different numbers. The whole invocation envelope lives inside
IToolInvocationService, so the executor enforces none of it and cannot skip any of it: every refusal arrives as an
outcome and becomes a row.
A landed outcome maps to a terminal: Executed succeeds; UnknownTool, NotInvocable and InvalidArguments fail
ValidationFailed and are therefore never re-attempted; Timeout and Faulted fail on the two retryable classes.
The service's own reason is repeated verbatim and is structural by contract — it never echoes an argument value. An
output document over the cap is a real, reachable outcome (a knowledge-base search may legitimately answer with fifty
thousand characters) and is deliberately not retryable, since the same call composes the same bytes; the message
names the tool beside the node, because "which node" alone does not say what to shrink.
graph-workflows/tools is the picker's feed, the same list the save gate checks (GraphWorkflowToolGate.ListAsync:
the invocation envelope, plus web_fetch flagged requiresAllowedUrls while web access is on), so the picker cannot
offer a name the run would then refuse. A missing binding path fails the node
ValidationFailed; the Queued write happens before bindings resolve, so a refused row keeps its input document.
Config: path — the optional node-level default dot path its own out-edges inherit.
Output: a verbatim pass-through of the predecessor's output. This is what makes a Condition node a real router:
edge conditions evaluate against the source node's output document, so without the pass-through a Condition's own
out-edges would inspect {} and never fire. A node with several predecessors has no single upstream output to carry
forward and answers {} rather than inventing one.
No config — they are shape, not settings, and §2.3 is the whole of their semantics.
Parallel passes its predecessor's output through exactly as Condition does. Join answers a per-source map
over its satisfied inbound edges, { "<nodeKey>": <that node's envelope> }, so everything downstream of a join
sees every branch rather than whichever one the single-predecessor shortcut would have picked.
| Config member | Meaning |
|---|---|
prompt |
Required. What the person is being asked. |
allowedDecisions |
Required, non-empty, distinct, from Approve / Reject. Empty would be a question nobody could answer. Answer is refused: it belongs to ChatInput (§4.8). |
requireComment |
Defaults false. |
GraphWorkflowPauseExecutor is the one lane that drives nothing: parking a row on a human is two status writes, so
there is no work to hold, no slot to wait for and no answer to poll for. It writes Running then
WaitingForApproval with PendingDecisionKind naming the pending act, and no output document — a pause's output is
its answer, and it has none yet. The prompt, the allowed answers and requireComment are not copied onto the
row: they are already in the pinned graph, and a second copy is a second thing that can drift.
The answer arrives through POST .../runs/{runId}/nodes/{nodeKey}/decide, idempotent on a client-minted
operationId. Both answers succeed the node run — the answer is the node's output, and routing on it is the
edges' job. A rejection reaches the run through an out-edge, never through a node failure. WaitingForApproval can
never move to Skipped: skipping an open pause would be an operator walking past a decision instead of giving one,
which is the one thing a pause exists to make impossible.
Output: { "decision": "Approve" | "Reject", "comment": … | null, "payload": … }. decision is the enum's name,
because that is the member every out-edge condition selects on and the exact member the definition-time pre-flight
check writes; the two spellings must produce the same string or a graph that pre-flighted clean would route nowhere.
comment and payload ride beside it rather than above it: they are the operator's, not the router's, and a
condition that could select on them would route on free text.
A Pause's successor gets only the decision, and the editor wires around it. A node's input is its ONE
satisfied predecessor's output (§3.2), so an authored X → Pause → Y hands Y the approval and never X's answer, and a
Pause before an End loses the result the same way. Two things address that, neither of them a change to what a
Pause writes. The validator raises the non-blocking warning of §2.4 on Y. And the editor adds the missing edge itself:
when an author connects an edge into or out of a Pause, pauseContextEdges returns the unconditional edges
labelled context from the pause's nearest non-Pause ancestor to its successors, the canvas adds them, and a notice
says one appeared. It descends from the Open Canvas importer's CanvasWorkflowImport.AddPauseContextEdges
(§9.2) — walking back through consecutive pauses, skipping a self-loop — judged per successor exactly as the
validator judges it: only a successor whose every inbound edge leaves a Pause is starved (one something else already
feeds is not, so it gets nothing and no warning), and the ancestor is the union of the nearest non-Pause ancestors
of every pause feeding it. The importer applies the same rule and the same guards (§9.2), so the three halves cannot
advise different edges. A Condition ancestor is skipped, because the added edge carries no
sourceHandle and would save as that Condition's second unconditional out-edge, which §2.4 refuses. A pause whose
nearest non-Pause ancestor is not unique (mutually exclusive branches feeding it) gets nothing, because wiring
both ancestors into an All successor would skip it the moment the untaken branch is dead. And a successor with
joinPolicy: "Any" — of any kind, since the policy is a property of every node (§2.3) — gets nothing, because an
unconditional content edge would admit it while every approval was rejected. Wiring into a pause visits each pause
reachable forward through consecutive pauses with its own ancestry, so A → P1 → P2 → B plus X → P2 adds nothing
for B (its ancestors through P2 are {A, X}). Everywhere else Y keeps its default All join policy, so it is admitted only once both
the content and the approval have arrived, and its input is the upstream map carrying both.
The pass runs on the connect gesture and nowhere else — never on render, never on validate — so an edge the
author deletes stays deleted. Wiring out of a Pause considers only the node just connected, since that is the
only successor that gesture can have starved; wiring into a Pause considers every successor of it, walking forward
through consecutive pauses so A → P1 wired last still reaches the B behind P1 → P2 → B.
A decide call carrying a different operationId for an already-answered pause is 409
GraphWorkflowGateAlreadyDecided, with the standing decision on the body — a second human act is refused, not
replayed.
| Config member | Meaning |
|---|---|
outcome |
Required. The declared outcome string. |
resultPath |
Optional dot path into the End node's input document. |
publishToChat |
Defaults true in a Chat graph, false otherwise; true is refused outside a Chat graph (§2.1). |
Output: { "outcome": …, "result": … } — the resolved path, or the whole input document when the author named none.
A path the document does not carry resolves to null; failing the node instead would end a run that did all of its
work over a projection nobody reads.
An End node is a terminal node by construction (nothing may leave it), and a run is Completed only once one of
them succeeded.
| Config member | Meaning |
|---|---|
prompt |
Required. What the chat surface shows while the run waits for the user's next message. |
Legal only in a Chat graph, and it needs at least one unconditional out-edge (§2.4). It rides the pause lane:
GraphWorkflowPauseExecutor owns Pause and ChatInput alike and writes PendingDecisionKind from the node kind —
Approve for a pause, Answer for a chat input — so a reader tells a gate from a question off the row, without
the graph. The run reads WaitingForApproval (the wording is the surface's, derived from the pending kind); a restart
leaves the row alone exactly as it leaves a pause, and the cancel drain cancels it the same way.
The answer arrives through the same decide route with decision: "Answer" and payload: { "text": "…" }.
GraphWorkflowStateMachine.IsDecidable(kind, status, decision) is the one place the pairing lives: a ChatInput
takes Answer and nothing else (409 for Approve / Reject), and a Pause refuses Answer (409). A missing, blank
or non-string payload.text, or a comment beside it, is a 400; so is text over
min(Security:MaxMessageSizeKb × 1024, MaxOutputJsonBytes / 2) UTF-8 bytes — the answer is a chat message, and it must
still fit the envelope it is embedded in. operationId idempotency and the 409 GraphWorkflowGateAlreadyDecided
naming the standing answer work exactly as for a pause.
Output (GraphWorkflowDocuments.ChatInputOutput): { "decision": "Answer", "text": … }. decision is kept because
the replay and standing-conflict checks read output.decision structurally. An edge may condition on output.text.
The §2.4 pause-context warning and the editor's pauseContextEdges gesture treat a ChatInput like a Pause: a
successor reached only through one receives the answer document, not the content before the wait.
| Config member | Meaning |
|---|---|
question |
Required. What is being decided. |
labels |
Required. 2–32 distinct, non-blank strings of at most 64 characters. Flat on purpose: they become a grammar enum, far below the repetition bound. |
provider |
Optional closed vocabulary, llm only (and the default). A provider this build cannot run is refused at save. |
model |
Optional, as on LlmCall: an installed node-managed GGUF chat model, else the local default. |
inputBindings |
Optional, exactly as on LlmCall (§4.2a). |
A classifier node on the invocation lane. The named IGraphWorkflowDecisionProvider
(Services/GraphWorkflows/Decisions/) lowers the node to an LlmCall config — a fixed classifier system prompt,
the question plus the label list as the prompt, the same bindings, reasoningEffort: "none", temperature: 0 (so the
same input routes the same way on a re-run) and the response schema
{ type: object, properties: { choice: { type: string, enum: labels } }, required: [choice] } — which runs through
RunLlmTurnAsync unchanged, model gate, capacity and binding rules included. The provider then interprets the
answer, and the executor holds the choice to the labels: an answer that is not a JSON object, or names anything
outside them, fails NodeFailed (retryable, so AttemptsExhausted on the last attempt), and the reason never repeats
the model's text. LlmDecisionProvider is the one provider; Laya/ONNX classifiers are later providers of the same
seam.
Output: { "choice": "<label>", "confidence": null, "probabilities": null, "provider": "llm", "usage": { … } }.
confidence and probabilities are always null from the llm provider — a grammar-constrained answer carries no
calibrated score, and inventing one would route on noise. Out-edges route on output.choice, directly or through a
Condition. An interrupted Running DecisionModel row is failed Interrupted at startup like every model turn (§3.5).
Full route table and hub inventory: API & Hubs. LocalApiRoutes.GraphWorkflows is the family;
the whole surface, hub path included, sits behind request-path middleware in Program.cs that answers 404 when
GraphWorkflows:Enabled is false — ahead of the security middleware, so the switch cannot be probed by status code.
Every route is Operator-gated.
| Route | Notes |
|---|---|
graph-workflows/definitions |
GET lists without the graph blob (it is the encrypted column), each row carrying its denormalised kind; POST creates. |
graph-workflows/definitions/{definitionId} |
GET / PUT with the version it was edited from / DELETE, which 409s while a live run pins the definition. |
graph-workflows/definitions/validate |
POST a graph and get its errors and warnings back without saving. The editor asks the runtime's own parser. valid is still zero ERRORS: a graph that only warns passes here and saves. |
graph-workflows/tools |
The Tool node picker's feed (§4.3). |
graph-workflows/definitions/{definitionId}/runs |
POST start → 202 with the run id; requestId is the idempotency key. |
graph-workflows/runs |
The run list, newest first, ?status=&limit= with limit required and capped at 200. |
graph-workflows/runs/{runId} |
One run, its node-run summaries, the run's own resolved output, and the graph this run pinned at start. No node-run documents — those are a per-node read. |
graph-workflows/runs/{runId}/cancel |
202. Live node runs drain first, so the run reads Cancelling. A repeat cancel is an idempotent 202. |
graph-workflows/runs/{runId}/nodes/{nodeKey} |
One node run in full, input and output documents included. |
graph-workflows/runs/{runId}/nodes/{nodeKey}/decide |
Answers a pause (§4.6) or a chat input (§4.8). |
graph-workflows/runs/{runId}/nodes/{nodeKey}/steer |
POST { operationId, message } → 202 with the run detail. Steers the queued or running Agent/LLM Call node of a chat-bound run (§3.7). 400 for the request, 404 for an unknown run or node, 409 GraphWorkflowSteerLimitReached at the cap, 409 GraphWorkflowRunConflict for the run's state or a reused id. The node-run read carries steering: [{ operationId, message, atUtc, attempt, applied }]. |
graph-workflows/runs/{runId}/events |
The event log, paged from an exclusive afterSeq, capped at EventReplayLimit — which the response reports rather than leaving a client to infer it from a full page. |
graph-workflows/conversations/{conversationId}/messages |
POST a chat send into workflow mode → 202 { runId, messageId, action } (§3.6, Chat). 409s: GraphWorkflowRunBusy, GraphWorkflowRerunConfirmationRequired, GraphWorkflowAttachmentsNotAccepted. |
graph-workflows/conversations/{conversationId}/runs |
GET the conversation's bound runs newest first (?limit=, default 20): run summary, definitionId, definitionName (null once deleted), triggerMessageId, pendingInput, steerable. |
Six routes cap the request body at 1 MiB (GraphWorkflowRequestSizeLimit): create, update and validate, which
carry a graph, start-run, which carries an input document, the chat send and the steer. Without it they would inherit Kestrel's 30 MB default
and a body that size would be parsed, walked and hashed before the node cap could refuse it. Kestrel enforces the cap
as it reads, inside model binding, so the 413 comes from RequestBodyTooLargeExceptionHandler rather than from the
endpoint. A name is capped at 200 characters, a description at 1024.
The run detail carries the run's own graph. GraphWorkflowRunResponse is (Run, NodeRuns, Output, Graph).
Graph is the copy this run pinned when it started, not the definition's current one, in the same wire shape
GraphWorkflowDefinitionResponse.Graph carries — so a client parses a definition and a run with one piece of code.
It sits on the run DETAIL and deliberately not on the run summary: a list must not carry one graph per row, and a
graph may hold up to a mebibyte. output is the run's resolved result, written once by the succeeded End node at
terminalization; the per-node input and output documents remain a separate read. Without the pinned graph a run view
drew the definition it names, which is the wrong graph for every run started before an edit — the whole reason a run
pins a copy at all.
A pinned graph that will not deserialize throws, where a node-run document reads as null. The difference is what
each blob is: a node-run document is written by the runtime and a broken one is worth reading a page about, while a
graph was parsed before it was ever stored, so no supported route reaches a corrupt one. An empty canvas drawn beside
a real nodeCount would report the corruption as a graph nobody drew.
Validation errors are keyed. GraphWorkflowValidationException carries a GraphWorkflowValidationResult — a list of
(key, message) pairs where the key is the node or edge, or null for a failure about the document as a whole — and
the endpoints replay them one by one instead of collapsing them into a sentence. It is therefore kept out of
DomainValidationExceptionHandler, which maps single-message validation exceptions globally.
Warnings travel on the validate response as a second list of the same (key, message) shape. They are a separate
Warnings member on GraphWorkflowValidationResult — never mixed into Errors, which is what the exception path and
valid read — so nothing that refuses on the errors has to filter them out, and a client that ignores the member
behaves exactly as it did before it existed.
GraphWorkflowRunHub at /api/local/v1/graph-workflows/hub. SubscribeRun(runId, afterSeq) joins the per-run group
graph-workflow-run-{runId:N} before reading the replay, so a change published between the read and the join
cannot reach nobody; the overlap that creates is harmless, because every push is idempotent and keyed by sequence.
The snapshot carries the run status, the queued / running counts, pendingDecisions (parked rows whose
PendingDecisionKind is not Answer — a pause) and pendingInputs (parked rows waiting on an Answer — a chat
input), the watermark, up to
EventReplayLimit events, and a replayTruncated flag read from one row past the limit rather than inferred from a
full page. There is no in-memory buffer — the store is the replay authority — and a disconnect cancels nothing,
because a run outlives the tab. UnsubscribeRun(runId) leaves the group, and StreamNodeActivity(runId, nodeKey)
streams one running node's live turn (below).
The pushed event is graphWorkflowChanged, carrying (runId, seq, kind) and no content at all: the subscriber
re-reads the named feed from its own watermark, so a dropped push degrades to a late read rather than to a wrong
render. kind is lowercase on the wire — run, node, gate — written as literals in
GraphWorkflowEventPublisher.ToWireKind and asserted literally on both sides. There is no event kind: every kind
moves the append-only event feed, so the client invalidates it unconditionally.
The publisher is a store decorator (PublishingGraphWorkflowStore), so exactly one ping is emitted per committed
mutation that has a subscriber, carrying the sequence that commit allocated — and every published transition carries a
fresh, increasing sequence. Three writes deliberately publish nothing, because nothing is subscribed to them: the
definition writes (CreateDefinitionAsync, UpdateDefinitionAsync, DeleteDefinitionAsync), StartRunAsync, and
the startup reconciler ReconcileNonTerminalNodeRunsAsync. API & Hubs states the same. A publisher failure is logged and never fails the write that already committed. Client.Application
depends only on IGraphWorkflowEventPublisher; the host swaps the hub-backed implementation in over a registered
no-op, so a host without the hub stays resolvable.
The event vocabulary is the closed nineteen-token GraphWorkflowEventTypes catalog — run.created, run.started,
run.waiting, run.completed, run.failed, run.cancelled, node.queued, node.started, node.completed,
node.failed, node.skipped, node.cancelled, node.interrupted, node.retried, gate.requested,
gate.decided, node.published (amendment 2026-09-23, Chat Workflows S1: a result became a chat message, §3.6),
and node.steered / node.steer-ignored (amendment 2026-09-23, Chat Workflows S3: a steer reset its row, or arrived
after the row settled, §3.7). The feed is append-only and durable, so a token written once is a token every later reader must
understand: extend it by amendment, never silently. Event details are small structured payloads — a failure summary,
a decision outcome — and never a transcript.
Live activity (Chat Workflows S2). StreamNodeActivity(runId, nodeKey) is a server stream over the node run's
InvocationId, served straight from IInvocationResumeRegistry.ResumeAsync: a snapshot, then offset deltas for
content, reasoning, tool calls and phases, then the terminal event — the same ChatStreamEvents
LocalChatHub.ResumeMessage serves, so the client reuses the chat stream folding. The registry already mirrors every
invocation from IWorkerEventDispatcher.InvocationStateChanged, graph turns included, so there is no new publisher
and no content-bearing group: graphWorkflowChanged stays content-free. The stream is not tracked as a chat
attachment (LocalChatHub.TrackAttachment), which would mark the graph turn detached for the reaper. It refuses with a
HubException when the run is unknown, the node key is not a row of this run, the row is not Running, it has no
InvocationId, or the registry no longer holds the turn; a client re-subscribes when a node ping changes the
running row's InvocationId. The registry is in memory, so a mid-node restart loses the live text: the row is failed
Interrupted and retried by the existing rule (§3.5), and the node's durable output is what remains.
Client.React/src/features/graphWorkflows/ — see React Client for where it sits among the
feature folders. One route, /graph-workflows, with all four selections (definitionId, runId, nodeKey, tab)
as search params rather than path segments, so every view is linkable and a reload lands back on it. The route
component is a thin adapter; GraphWorkflowsPage itself is router-free and is rendered directly in unit tests.
| Directory | Holds |
|---|---|
models/ |
GraphWorkflowModels.ts (the one file naming generated DTOs, plus the closed vocabularies and narrowers), GraphWorkflowCanvasModels.ts (the discriminated node union and the graphToCanvas / canvasToGraph round trip), GraphWorkflowLayout.ts, GraphWorkflowValidation.ts (the client mirror of the graph rules), GraphWorkflowRunGraph.ts. |
queries/ |
Every read and mutation over the generated adapters, including the forward-paged events feed, the several-node-runs read (useSettledGraphWorkflowNodeRuns) and the two chat routes (useGraphWorkflowConversationRuns, useSendGraphWorkflowChatMessage). |
hooks/ |
useGraphWorkflowEditor (controlled React Flow state, per-handle connect prefill, refusal of a second unconditional edge, and the context edge added around a Pause or ChatInput on connect — §4.6) and useGraphWorkflowRunHub. |
components/ |
Editor: the per-kind node cards, the canvas with its palette and Auto-arrange, the validation strip, the node and edge config panels (the DecisionModel body and the input-bindings list shared with LlmCall under config/), the workflow settings popover (graph kind and the chat block, saved with the graph), the definition list and meta dialog. Run view: the status badge, the read-only run graph, the node-run table, the run list, the events tab, the node panel and the decision panel (Approve/Reject buttons for a Pause, a text-answer form for a ChatInput). |
samples/ |
The sample graphs the New workflow dialog can start from, and their index GraphWorkflowSamples.ts (see Samples below). |
pages/ |
GraphWorkflowsPage — editor mode without a runId, run mode with one. |
api/ |
GraphWorkflowConflict.ts, which reads the NodeConflictProblemType members by name — the three run/definition/gate refusals and the three chat-send ones (GraphWorkflowRunBusy, GraphWorkflowRerunConfirmationRequired, GraphWorkflowAttachmentsNotAccepted). |
features/chat/workflow/ (the chat side, not this folder) |
useChatWorkflow (the chat page's workflow mode: picker state, send routing, 409 handling, the live bound run's hub subscription), WorkflowSelectorCard, WorkflowRunStatusCard, WorkflowActivityBlock, ChatWorkflowModels.ts (taken path by layout rank, activity rows), ChatWorkflowStore (UI state only). It imports this feature's queries, hub hook, layout, models, status badge and conflict reader — ten reviewed no-cross-feature fingerprints in config/dependency-baseline.json. See Chat. |
The editor mirrors the server's rules rather than inventing its own, and the mirror is a mirror on purpose: the
authoritative answer comes from graph-workflows/definitions/validate, which runs the parser a run would run, and
serverErrorsToIssues maps its keyed errors back onto the canvas elements. The page flow is Validate → Save (server
validate first; a 409 reloads) → Start. The server's warnings arrive through serverWarningsToIssues as
issues with the rule serverWarned and reach the strip through its own warnings prop; they render in their own alert
below the errors, in a different colour and under a title that asks rather than refuses, and never join the list the
save gate, the red chips and the config panels read, because a warning is not a reason to save less. Every client rule mirrors a rule that REFUSES a save, so the client raises none
of them.
Two round-trip rules the editor has to hold because the server does. Every JSON-shaped config field is written back as
JSON text with a string quoted, since that text is exactly what a save parses again — an unquoted string read as
the invalid-JSON complaint on a defaultInput the server had accepted. And the client's dot-path mirror refuses an
EMPTY segment the way GraphWorkflowTokens.IsDotPath does, so a..b cannot read green in the drawer and then 400
on save.
Samples. The New workflow dialog offers "Start from" with a blank workflow or one of the samples in
features/graphWorkflows/samples/: one wire graph per <id>.json, indexed by GraphWorkflowSamples.ts with an i18n
name and description. Picking one prefills the name and description (never over text the operator typed), and the
create call posts the sample's graph instead of the Start → End starter, so a created sample is an ordinary definition
with no link back to its file. Every sample must run on any installed node-managed GGUF chat model and nothing else:
no Tool or ChatInput node, no model pin on any node, a Standard graph, and a Start defaultInput so the Start
Run dialog is one click. An Agent node is allowed only without agentDefinitionId: unbound, it runs the built-in
default persona with the local tools a fresh install already has (§4.2), which is how agent-research-and-check shows
where an agent fits (Agent → LLM Call check → Pause → End); a bound agent would need setup first. Prompts inside the graphs are model-facing config and stay English. GraphWorkflowSamples.test.ts
holds the client half (validator, lossless canvas round trip, bundle keys), and
XE-Local-AI-Engine.Tests/GraphWorkflows/GraphWorkflowSampleContractTests.cs parses every file through the real
GraphWorkflowGraph.Parse and also refuses any warning, so a sample never opens with a note the operator did not cause.
useUnsavedChangesGuard is set with allowSameRoute, so writing a search param — selecting a node or a tab — does
not trigger the leave prompt, while a real route change still does.
Nodes are laid out left-to-right by layoutGraphWorkflow, a dependency-free layered DAG layout (longest-path ranking
in Kahn order, one barycenter pass, ties broken by node key). It is used for three things: a node that arrives
without a position, the editor's Auto-arrange, and the run view's nodes-only fallback for a response that carries
no graph at all. A definition whose nodes carry no positions opens laid out and dirty — the layout is a proposal
the operator has not saved yet. That is by design, and §9 is why it matters.
The run view is read-only, and it draws the graph the run pinned at start — the one GET runs/{runId}
carries (§5) — with each node run's state overlaid on it. A definition edited since the run started is then worth
saying and nothing more: an informational notice that this is the graph the run itself ran on, never a reason to
draw less. Nodes only, auto-laid-out, is what remains for a response that carries no graph and whose hash
disagrees with the definition on screen, because drawing today's edges over an older run would be a lie about routing.
The Pause decision panel reads its prompt, its allowed answers and requireComment off the pinned graph for the same
reason: a definition edited since could offer a decision this run's gate would refuse. The view also lists the node
runs, shows a selected node's input, output and error documents, and renders the event trail. The events feed pages
forward from afterSeq: 0, so replayTruncated means the newest events are missing, and the banner says
exactly that.
The graph → canvas conversion is memoised on the graph object. graphToCanvas parses the whole document and lays
out every node that carries no stored position, while the run hub invalidates the run detail on every event — so one
node transition would otherwise re-parse and re-rank up to a mebibyte of graph to redraw a single badge. The cache is
a WeakMap keyed on the graph object itself: the conversion is a pure function of that document, TanStack's
structural sharing hands back the same object while its JSON is unchanged, and an entry dies with the graph that
keyed it. Callers treat the cached canvas as frozen and copy every node and edge they annotate.
useGraphWorkflowRunHub degrades rather than fails. On a subscribe it re-sends its watermark as afterSeq; the
store serves a client that has been away for days, so there is no buffer to roll over. While the hub is unavailable
it hands the page a poll interval of 3 s, cleared only on a successful subscribe. Node-run invalidation is keyed
on the run, not on the selected node, so clicking through the node table does not tear down the subscription.
The feature ships on and is a top-level navigation entry. Its capability flag, nodeCapabilities.graphWorkflows,
is declared in src/capabilities/NodeCapabilities.ts with the rest; the route redirects home when it is off.
See Writing Tests for which project a new test belongs in and Testing & Validation for how to run them.
| Layer | Where |
|---|---|
| Parser, conditions, documents, state machine, deadline, dispatcher, executors, run service, options | XE-Local-AI-Engine.Tests/GraphWorkflows/ |
| Endpoints | XE-Local-AI-Engine.Tests/Endpoints/GraphWorkflows/V1/ |
| Hub and event publisher | XE-Local-AI-Engine.Tests/Hubs/GraphWorkflow*Tests.cs |
| Store, encryption, purge coverage | XE-Local-AI-Engine.Client.Persistence.Tests/GraphWorkflows/ |
| React components, hooks, models, queries | colocated *.test.ts(x) under features/graphWorkflows/ |
| Browser end-to-end | XE-Local-AI-Engine.Tests.E2ETests/Tests/GraphWorkflowE2ETests.cs |
Two suites are worth naming because they are the ones that catch a wiring break the unit tests cannot:
GraphWorkflowEndToEndTestsandGraphWorkflowPauseToolEndToEndTestsdrive the wired host over REST and a real SignalRHubConnection, with the model runtime replaced byTesting.FakeOllama. The pause/tool suite asserts thegatepush twice, which is what proves the publisher decorator and the lowercase wire kind agree with the client.- The E2E class drives the real browser through the whole loop — create a definition, drop nodes from the palette, wire and configure them, validate, save, start a run, answer the pause, read the event tab.
GraphWorkflowHarness and the *HostFixture types are the shared seams; GraphWorkflowGraphs holds the reusable
graphs. On the client, test/GraphWorkflowFixtures.ts is the one fixture file, and eightNodeGraph is the graph of
§2.5.
GraphWorkflowOptions, bound from the GraphWorkflows configuration section. No appsettings*.json carries the
section — the defaults below are the shipped values, and an operator overrides one through configuration or an
environment variable such as GraphWorkflows__MaxConcurrentRuns.
| Option | Default | What it bounds |
|---|---|---|
Enabled |
true |
The whole surface. False makes every route and the hub answer 404 through the request-path middleware, while registration stays intact so a disabled node answers legibly instead of 500-ing out of an empty container. |
MaxNodesPerDefinition |
200 | Nodes in one definition, enforced when it is validated rather than when it runs, and checked against the declared node count before the graph is parsed (§2.4). A run that re-parses an already-stored graph is deliberately uncapped. |
MaxNodeRunsPerRun |
200 | Node runs one run may instantiate. Never below MaxNodesPerDefinition — a run that could not instantiate the definition it started from would fail halfway through a graph the operator was allowed to save. |
MaxTotalAttempts |
50 | Every attempt one run may spend across all its nodes. The guard against a retry storm. |
DefaultNodeTimeoutSeconds |
600 | One node run's attempt, when its node names no timeoutSeconds. Unlike Dev Workflows, a node that declares nothing still has a deadline. |
MaxOutputJsonBytes |
262 144 | One node run's composed output document, in UTF-8 bytes, checked before it is encrypted and stored. |
DispatchIntervalMilliseconds |
500 | The sweep cadence, independent of the change signals the dispatcher also listens for. Floored at 100 ms. |
MaxConcurrentRuns |
4 | Executing runs at once — a run parked on a person holds no slot (§3.4) — and the size of both the shared Agent/LLM Call/DecisionModel invocation lane and the Tool lane. Runs above the cap wait; they are not refused. |
MaxRunInputBytes |
65 536 | A run-start input document, checked in GraphWorkflowRunService.StartAsync (§3.1) rather than at the endpoint, so every caller of the service is held to it. Also the budget for the inlined upstream map in an Agent prompt and the complete user prompt with bound JSON data in an LLM Call. |
EventReplayLimit |
200 | Events one replay may return, hub snapshot and events route alike. Ceiling 1000 — one replay is one response body. |
MaxSteersPerNode |
5 | Steers one node run may take (§3.7). Each applied steer re-runs the node on the same attempt and restarts its deadline, so the attempt budget does not bound steering; this does. Floored at 1. |
GraphWorkflowOptionsValidator checks at startup what the data annotations cannot: a semantic floor under each
budget (a MaxNodesPerDefinition of 1 passes [Range(1, …)] and still admits no graph, since every graph carries a
Start and an End), the MaxNodeRunsPerRun ≥ MaxNodesPerDefinition relation, and the replay ceiling. An operator meets
these at boot rather than once per node run.
Two limits that are not options, because they are not runtime budgets: the 1 MiB request-body cap on the five routes that carry one (create, update, validate, start-run and the chat send), and the 200-row cap on a run list page.
One ceiling that lives outside this module: an Agent node's responseJsonSchema goes down the same llama.cpp GBNF
path as a tool schema, which has an empirical combined repetition bound
(LlamaGrammarToolSchemaCompatibility.MaxGrammarRepetitionBound, 1024). Keep a response schema flat. Since S7 that
sanitiser covers the response schema too: DeferredLlamaServerChatClient.ApplyResponseSchemaPassthrough runs the
authored schema through the same Sanitize pass before it reaches the wire, so an over-large bound is dropped rather
than failing the turn with HTTP 400 Failed to initialize samplers. Every bound within the cap is now enforced by
the grammar rather than dropped — see §4.2. See
Local Runtime & Providers and docs/agent-knowledge.md §2.
Open Canvas (the Preview workflow builder) was removed. Its saved canvases are converted into Graph Workflow definitions automatically during startup. There is no button, prompt or opt-out. The encrypted source survives interrupted conversion; startup stops on conversion failures and retries after the underlying fault is resolved.
Back up the node's data directory before upgrading a node that has canvases you care about. That is not a formality: see the failure posture below.
An EF migration cannot decrypt graph_json — it has no node key, and the column is AEAD ciphertext — so the
conversion has to run in application code. But the DropCanvasWorkflows migration removes canvas_workflows during
ApplyNodeChatMigrationsAsync, which runs before any hosted service starts. A single post-migration importer would
therefore always find the table gone.
So the step is split, in Program.cs:
var pending = await ReadPendingCanvasWorkflowsAsync(app.Services); // read + decrypt, BEFORE migrations
await ApplyNodeChatMigrationsAsync(app.Services); // backs up, then runs DropCanvasWorkflows
await ImportCanvasWorkflowsAsync(app.Services, pending); // write, IMMEDIATELY after that pass
await ApplyNodeIdentityMigrationsAsync(app.Services); // a different database; runs after the write
Before migrations, the reader uses one SQLite transaction to copy the legacy rows into
canvas_workflow_import_recovery without decrypting or rewriting their graph bytes. This is a durable staging
table in the node database, not a SQLite temporary table or a restored Open Canvas feature. While the original table
exists it remains authoritative, and a retry atomically refreshes staging from it. Once migrations drop the original,
startup reads staging instead. A process interruption between migration and import therefore loses no source rows.
After migrations, the importer uses the same scoped NodeChatDbContext as the definition service and store. It writes
all definitions and drops the recovery table in one transaction. A write, cleanup or commit failure rolls back
both operations; a later startup imports the retained ciphertext without duplicate definitions. Fresh installs have
neither table and import nothing. An empty legacy table is staged and cleaned up without creating definitions.
The import runs regardless of GraphWorkflows:Enabled, deliberately: an operator who never turns the feature on must
not silently lose their canvases. Every row is read — there is no cap — because a cap plus an unconditional drop
in the same build would destroy everything past it as its normal outcome.
Nothing in the importer may depend on a Preview type: the read parses the stored blob into private records of its own,
so the Open Canvas namespace can be deleted out from under it. It also runs under --reset-admin-password, whose
branch of Program returns only after the migration pass, so the read, the migrations and the write all happen
first. The knowledge-downgrade commands are the one launch path that returns before migrations, and therefore before
the import.
| Open Canvas node | Graph Workflow node |
|---|---|
Start |
Start, with the canvas StartText as config.defaultInput = { "text": <StartText> } (null when empty) — an object, because the editor renders a stored default as JSON text and a bare string would never re-parse. |
Agent |
Agent with maxAttempts: 1, instructions / model / reasoningEffort carried over, agentDefinitionId: null, responseJsonSchema: null, includeUpstreamOutputs: true. |
Debug |
Elided. Every X → Debug and Debug → Y collapses to X → Y. A Debug node was a side-event tap that forwarded its input unchanged, so removing it preserves the run's meaning exactly. |
Pause |
Pause with prompt: "Approve and continue?", allowedDecisions: ["Approve"], requireComment: false; its single out-edge gains label: "approved" and condition: { "path": "output.decision", "op": "Eq", "value": "Approve" }. Open Canvas's Continue was a resume rather than a decision, so Approve alone is the faithful translation — and one allowed decision with one matching out-edge satisfies the Pause pre-flight rule of §2.4. Plus one context edge X → Y, unconditional and labelled context, where X is the pause's unique nearest non-Pause ancestor along the (Debug-elided) chain and Y each starved successor of it — see below. |
End |
End with outcome: "completed", resultPath: null. |
ModelProfile (on an Agent) |
Dropped. An Agent node's config has no profile member. A non-null value records a reason naming the canvas id and the node key, never the value. That reason reaches the log as a per-change Warning only when the definition imported cleanly; on the needs-attention path it is folded into the description instead (§9.3). |
Node keys are the Open Canvas ids sanitized to [A-Za-z0-9_-]{1,64}; an unmappable or duplicate key becomes
n{index}, and a sourceId → key map rewires the edges. Each emitted edge gets a minted key e{index}. Node and
edge keys share one namespace, so CanvasWorkflowImport.MintKey appends _2, _3 and on when the fallback is
itself taken — the key is unique before it is pretty.
The definition takes the canvas name, truncated to 255 characters (CanvasWorkflowImport.MaxNameLength, which
is the name column's own bound — GraphWorkflowDefinitionConfiguration declares HasMaxLength(255)), and the
description "Imported from Open Canvas (canvas workflow {id})." — provenance without a schema change. The create and
update validators cap a name at 200 (GraphWorkflowRequestLimits.MaxNameLength), so an imported name of 201 to
255 characters lands in the database but has to be shortened before the definition can be saved again.
The importer emits no positions. position is optional and the editor already lays out a node that arrives
without one, so generating a second layout rule here would leave one of the two dead. The practical consequence:
an imported definition opens laid out and unsaved. The canvas is dirty the moment you open it, and the layout is
persisted when you save. That is expected, not a bug.
The context edge around a Pause. Open Canvas's Pause was a pass-through resume: its post-adapter forwarded
the answer it was waiting on unchanged. A Graph Workflow Pause writes a decision document of its own —
{decision, comment, payload} — and a node's input is its ONE satisfied predecessor's output document, becoming
the upstream map only when there are several (§3.2). Mapped one-for-one, Agent A → Pause → Agent B therefore
hands B the approval metadata and never A's answer, and a Pause before an End loses the result the same way.
That was observed live in the S4 round, not reasoned about.
So the importer emits one extra unconditional edge from the pause's nearest non-Pause ancestor X to each
starved successor Y. Y keeps the default All join policy, so it is admitted only once both the content
edge and the pause's own approved edge are satisfied — never ahead of the approval — and with two satisfied
predecessors its input is the upstream map { "<X>": …, "<P>": … }. An imported Agent carries
includeUpstreamOutputs: true and so sees X's text; an imported End carries resultPath: null and so keeps both
documents. The rule applies to every pause and walks back through consecutive ones, so A → P1 → P2 → B gains both
A → P2 and A → B; a successor several pauses reach is judged once, over the union of what all of them are fed
by. X may be the Start node, whose output is the run's own input, which is exactly the content the pause
interrupted.
The decision is keyed on the successor, and the guards are the validator's own
(GraphWorkflowGraph.PauseContextWarnings), read off the mapped node's joinPolicy and kind rather than off a canvas
kind — an importer that advised an edge the validator refuses would produce a definition nobody can save again. A
successor is owed an edge only when it is starved (every one of its inbound edges leaves a Pause; one something else
already feeds loses nothing), its joinPolicy is not Any (of any kind — an unconditional content edge would admit
an Any node ahead of every approval, including when all of them were rejected), and its pauses' nearest non-Pause
ancestor is unique and is not a Condition (two candidates are mutually exclusive branches, and an edge from each
would hang an All successor on the branch never taken; a Condition would gain a second unconditional out-edge,
which §2.4 refuses). Three more cases add nothing: a pause with no successor (already an IMPORT NEEDS ATTENTION:
graph, §9.3), a pair the canvas already wires — a second unconditional edge over one pair is a validation error
(§2.4) — and a self-loop.
Every guard but the starvation test is unreachable for a real import: an Open Canvas graph carries no Condition,
no Join and no join policy, and every stored shape is linear. They are stated in code anyway because the rule, not
today's canvas vocabulary, is what the next node kind has to keep holding.
The editor offers the same edge to an author, on the connect gesture (§4.6), under the same three guards. The one difference is when it runs: the editor runs only on that gesture, so an author who deletes the edge keeps it deleted, while the importer runs once over a whole canvas.
maxAttempts: 1 on an imported Agent is deliberately below the default of 3. An import is conservative: re-running
somebody's agent turn twice more, on a graph they have not looked at since it changed shape, is not a decision this
conversion gets to make for them.
The mapper is total — it always returns a document and never refuses a graph. Whatever it could not translate faithfully becomes a reason.
A converted graph then goes through IGraphWorkflowDefinitionService, which owns the parse, the node cap and the
hash and count. If that validation refuses it, the definition is saved anyway, through IGraphWorkflowStore and
past the validator, with its description prefixed:
IMPORT NEEDS ATTENTION: <reason>
Such a definition cannot be run until an operator opens it in the editor and fixes what the validation strip names. Nothing is discarded for being invalid. A Debug node with no successor, a shape the new validator refuses, an unknown kind — all of them arrive as a definition you can see and edit, rather than as an absence you have to notice.
A blob that does not decrypt or parse as JSON is reported at Error with its canvas id and exception type, then the read fails and startup stops before destructive migrations. The source and recovery ciphertext remain available for repair or restoration with the original node key. A whole-read or staging failure also stops startup; it is never treated as an empty successful read.
The write half records mapping or persistence failures at Error and refuses to commit any definitions when one row failed. Validation refusals still use the needs-attention path and preserve the graph. Only after every candidate has been preserved does the transaction remove recovery staging and commit. Startup propagates failures instead of serving a node whose conversion silently lost data.
Graph content is never logged. Instructions and start text are exactly the payload the column is encrypted to protect. The reader logs entry and per-row failures; the writer logs one line per clean import. The summary an operator must not miss is logged at Warning, only after the import and recovery cleanup commit:
Open Canvas one-shot import complete: {Imported} imported, {NeedsAttention} need attention, {Failed} failed.
Open Canvas has been removed; canvas_workflows is dropped.
One further Warning line follows per needs-attention canvas, naming the id, the name and the reason. A failed canvas gets an Error line instead, and it names no reason: the read-half line carries the id and the exception type, the write-half lines the id, the name and the exception type. A cleanly imported canvas whose mapping still changed something gets one Warning per change.
The same-database encrypted staging is the recovery mechanism for this conversion. The existing
INodeDbBackupService.BackupBeforeMigrationAsync() backup is supplementary: its best-effort failure no longer removes
the only recovery copy. If the durable staging write fails, startup stops before migrations. If import fails after the
source table has been dropped, staging remains and the next startup retries it.
Back up the complete data directory, including its key material, before upgrading. This fix cannot recover canvases that an earlier version already dropped without importing; those require a pre-upgrade backup. Do not delete or edit the recovery table to bypass a startup failure. Resolve the reported underlying fault or restore a complete backup. No plaintext export is created.
- Architecture Overview — where the module sits in the host.
- API & Hubs — the route family, the hub, and the event contract.
- Data & Persistence — the four tables, their encrypted columns and AAD binding.
- React Client — the
graphWorkflowsfeature among the others. - Agent Mode — the invocation stack an
Agentnode runs on, and the tool catalog aToolnode is filtered against. - Security & Privacy — the approval boundary the Tool gate tightens on.
- Testing & Validation and Writing Tests.
- ADR 0001 — the never-resume-a-stream restart posture §3.5 applies.
- ADR 0006 — the tool trust boundary §4.3 rests on.
{ "schemaVersion": 1, "kind": "Standard", // optional: "Standard" (default) | "Chat" "chat": { "acceptsAttachments": false, "requireRerunConfirmation": true }, // only with kind "Chat" "nodes": [ /* … */ ], "edges": [ /* … */ ] }