Skip to content

Latest commit

 

History

History
1373 lines (1125 loc) · 108 KB

File metadata and controls

1373 lines (1125 loc) · 108 KB

Graph Workflows — Operator-Authored DAGs

Reviewed: 2026-09-22 · Code-grounded.

Graph Workflows let an operator draw a directed acyclic graph of LLM calls, agent turns, tool calls, conditions and human pauses, save it, and start runs of it. The engine executes the run from the database: every node run is a row, every change is an append-only event, and the process that advances it holds no authoritative state of its own. Restarting the node loses at most the work that was in flight.

The module spans the whole stack: Client.Application/Services/GraphWorkflows/ (the parser, the state machine, the dispatcher and the node lanes), Client.Persistence (four tables, per-column AEAD encryption), the LocalApiRoutes.GraphWorkflows endpoint family plus GraphWorkflowRunHub, and Client.React/src/features/graphWorkflows/ (a React Flow editor and a read-only run view).

It replaced Open Canvas. The Preview / Open Canvas visual builder was removed in the same slice that turned this feature on by default; saved canvases are converted into Graph Workflow definitions once, automatically, on the first start of the new build. That conversion is §9; read its recovery and backup guidance before upgrading a node with canvases you care about.


1. What it is, and what it is not

What it is:

  • Operator-authored. A definition is a graph the operator drew. Nothing generates one, and no agent may edit one.
  • Manually started. A run begins because a person pressed Start (or a caller posted to the runs route with an idempotency key). There is no schedule, no trigger, no event subscription in v1.
  • Database-as-truth. GraphWorkflowDispatcher re-reads the rows on every tick and writes back through the store. Its only in-memory state is a parsed-graph cache keyed by run id, which exists to avoid decrypting and re-parsing the pinned blob on every tick and can be dropped at any moment without changing an answer.
  • Durable across a restart. A run outlives the browser tab and the engine. GraphWorkflowStartupReconciler makes the node runs a crashed host left in flight judgeable again, exactly once, before the dispatcher's pumps start.
  • Acyclic. GraphWorkflowGraph.EnsureAcyclic refuses a cycle at save time. There is no loop construct.

What it is not:

  • Not a general orchestration language. No boolean algebra in a condition (two comparisons are two edges into an All join), no expressions, no variables, no sub-workflows, no map over a collection.
  • Not MAF Workflows. Microsoft.Agents.AI.Workflows is not referenced anywhere in Client.Application; the routing here is this module's own state machine, and an Agent node reaches the model through the same headless invocation stack a scheduled saved-agent run uses.
  • Not a write surface for agents. A Tool node may only run a built-in read-local tool that needs no approval (§4.3). A node runs unattended, so there is nobody to ask.
  • Not Dev Workflows. Development Workflows run a fixed template over a code work item with gates, interventions and artifacts. Graph Workflows are free-form graphs with no work item and no artifacts, and several of their types are deliberate copies trimmed of what has no meaning here — there is no Blocked node-run state and no Waived edge state, because v1 has neither retry routing nor a waiving decision.

2. The graph contract

One JSON document per definition, stored encrypted as graph_json. GraphWorkflowGraph.Parse is the only entry point and parsing is the validation: a graph that survives it is one the dispatcher can route without a second opinion. Save time and run start call the same parser, so a graph accepted at save is a graph that will start.

{
  "schemaVersion": 1,
  "kind": "Standard",          // optional: "Standard" (default) | "Chat"
  "chat": { "acceptsAttachments": false, "requireRerunConfirmation": true },  // only with kind "Chat"
  "nodes": [ /* … */ ],
  "edges": [ /* … */ ]
}

schemaVersion is optional but, when present, must be 1. Anything else is refused outright.

kind says what the definition is for. A Chat graph is one the Chat page can bind a conversation to (Chat Workflows); it may carry ChatInput nodes and publishToChat flags. An unknown kind is refused outright, as is a chat block on a graph that is not Chat and any member of chat other than the two above. A Chat graph without the block reads the defaults shown. The kind is denormalised onto the definition row (graph_workflow_definitions.kind, plaintext, indexed, migration AddChatWorkflowNodes) at every save that carries a graph, and the definition list and detail report it, so a picker filters chat workflows without decrypting a blob. A Standard graph that names no kind is stored and returned exactly as before — the wire mapper omits both members when they are absent.

2.1 Nodes

Member Required Meaning
key yes 1–64 characters of letters, digits, _ and -. Node and edge keys share one namespace.
kind yes One of Start, LlmCall, Agent, Tool, Condition, Parallel, Join, Pause, End, ChatInput, DecisionModel, by NAME.
label no Display text. Defaults to the key.
joinPolicy no All (default) or Any. A property of every node — see §2.3.
maxAttempts no Positive. Defaults to 3 for LlmCall, Agent, Tool and DecisionModel, 1 for every other kind.
timeoutSeconds no Positive. Overrides DefaultNodeTimeoutSeconds for this node only.
position no { x, y }, both numeric. Authoring metadata the runtime never reads.
config no The per-kind settings, discriminated by kind — see §4.

config is closed per kind. GraphWorkflowGraph.ConfigMembers lists exactly what each kind reads, and a member no node of that kind reads is an author-time error rather than a setting that silently does nothing. Writing a Tool node's toolName on an Agent node fails the save.

Two chat members sit on several kinds and follow the same closed rule. publishToChat (Agent, LlmCall, End) marks a node whose answer a chat-bound run posts into the conversation; it defaults to true on an End of a Chat graph and false everywhere else, and true is refused outside a Chat graph, where no conversation could receive it. includeAttachments (Agent, LlmCall) hands the run's chat attachments to the node (§3.8), and true is refused unless the graph's chat.acceptsAttachments is on. Neither is read by routing; both are carried for the chat surface.

A node without a position is laid out client-side when the definition is opened (features/graphWorkflows/models/GraphWorkflowLayout.ts). That matters for §9: imported graphs carry no positions.

2.2 Edges and conditions

Member Required Meaning
key yes Same charset and namespace as a node key. Its identity — which is what makes parallel edges expressible.
from, to yes Node keys the graph declares. An endpoint the graph does not declare is a structural refusal.
label no The named outcome the source's output document reports as its branch.
condition no { path, op, value }. Absent means unconditional.
sourceHandle no Editor metadata. Read past and never stored.

A condition is a single declarative comparison against the source node's output document, evaluated by GraphWorkflowCondition.Evaluate. It never sees the target's input.

  • path is a dot path and nothing else: property names separated by ., with no wildcards, no array indexing and no functions (GraphWorkflowTokens.IsDotPath). items[0].name is refused at save rather than saved as a property literally called items[0].
  • op is one of Eq, Ne, Gt, Gte, Lt, Lte, Exists, NotExists, parsed by name (case-insensitively; a numeric token is refused, because Enum.TryParse would otherwise hand back an operator no member has).
  • value must be a scalar — string, number, boolean or null. An object or array has no comparison to make, and a relational operator against a boolean could never fire, so both are refused at authoring time.
  • Evaluation fails closed: a path the output does not carry answers false for every operator except NotExists. An edge must never fire on data that is not there.
  • Numbers compare exactly — long, then decimal, then BigInteger for integer tokens past decimal's range — so a Gt over ids past 2^53 answers on the numbers rather than on a rounding artefact. double is the fourth and last arm (GraphWorkflowCondition.Order), reached only by the fractional and exponent tokens no exact arm reads; two that differ only past roughly 17 significant digits therefore read as equal, which is the chain's stated ceiling. A token not even double reads is not an ordering at all. Strings compare ordinally. A type mismatch is not an ordering and reads as "no".

A Condition node may carry a default path in its own config; its out-edges inherit it when their condition omits one. That is authoring convenience only — the comparison still lives on the edge. An edge that resolves a path from neither is refused, because fail-closed would otherwise make it an edge that silently never fires.

Two edges over the same (from, to) pair are legal and are how an author widens a branch. At most one of them may be unconditional, since a second unconditional edge could only repeat the first.

2.3 joinPolicy is on every node

This is the trap the module documents most loudly, in GraphWorkflowEnums.cs, in GraphWorkflowStateMachine.Admission and again here: joinPolicy is a property of every node, not of Join nodes. An ordinary node with two inbound edges joins them exactly as a Join node does.

  • All (the default) waits for every inbound edge to be satisfied. One dead inbound edge skips the node.
  • Any proceeds on one satisfied edge, but only once no sibling could still satisfy one. The parser refuses Any with fewer than two inbound edges: one edge is an All written confusingly, and none would never fire.
  • All over zero inbound edges is vacuously satisfied, and that is load-bearing — it is how the Start node becomes eligible.

Pending outranks Dead under both policies. A dead inbound edge already settles what an All join will do, but settling it while a sibling branch is still running would skip the node, and everything after it, in front of work the run has not finished.

By the same rule, Parallel and Join are labels, not semantics. Fan-out is any node with more than one satisfied out-edge; fan-in is the join policy. The two kinds exist because they are inline nodes that write two rows each, which is what makes the timing of a fan-out visible in the event log at all.

2.4 Whole-graph rules

Structural failures throw immediately — there is nothing useful to say about the rest of a graph nobody can walk. Everything after them accumulates, keyed to the node or edge it belongs to, so an author fixing a canvas gets every complaint at once (GraphWorkflowValidationResult).

Throw-first:

  • exactly one Start node, with nothing routing into it;
  • at least one End node, or no run could ever complete;
  • acyclic;
  • every node reachable from Start.

Accumulated, per element:

  • a non-End node with no outbound edge (a run reaching it would stop without reaching an End);
  • an End node with an outbound edge;
  • joinPolicy: "Any" with fewer than two inbound edges;
  • a Condition node with fewer than two out-edges, or more than one unconditional out-edge;
  • a Pause node offering a decision no out-edge fires on — checked through the state machine's own routing over the document a pause actually produces, so the pre-flight rule and the routing cannot disagree;
  • a ChatInput node in a graph that is not Chat, or one with no unconditional out-edge (whatever the user types, the run must go somewhere);
  • a second unconditional edge between one pair of nodes.

Warnings are a second list, and they never refuse. GraphWorkflowGraph.Warnings is computed on the first ask rather than during the parse — only the validate endpoint asks, and every dispatcher tick parses — and nothing in it reaches GraphWorkflowValidationException. A graph that warns saves, validates as valid, and runs. There are two warning kinds in v1, and Warnings is their concatenation.

The first: a node whose inbound edges all leave a Pause receives the decision document rather than the content that was approved (§4.6). It is keyed on that node rather than on the pause, because that is the node which loses the content and so the node an editor draws the badge on, and one warning is raised however many pauses reach it. The sentence names the pause's nearest non-Pause ancestor only when that ancestor is unique and is not a Condition: two candidates means the pause is fed by mutually exclusive branches, and edges from both would leave an All successor waiting on the branch that was never taken, while a Condition cannot be named because the edge it would ask for is that node's second unconditional out-edge, which the parser refuses — advice that turns a warning into a hang or an error is worse than the generic sentence.

A successor whose joinPolicy is Any is exempt, whatever its kind, and that is a correctness rule rather than a taste one. The advised edge is unconditional, so it stays satisfied when every approval is rejected — an Any node would then be admitted on the content edge alone and run the branch the rejections were meant to stop. Collecting the decision documents is what such a node is for, so there is nothing to warn about. An All successor waits for the approval edges too, so it keeps both the warning and the advice.

The second, on an Agent node whose responseJsonSchema asks for something the runtime it will run on does not enforce (GraphWorkflowGraph.ResponseSchemaWarnings). It fires on three things: a keyword the Microsoft.Extensions.AI.OpenAI strict-schema transform relocates into description, a declared property missing from required, and an object that declares properties and omits additionalProperties — an object that sets that key explicitly, true included, is silent. The grammar is then built from the rewritten schema, so maxLength: 3 survives only as a hint the model may read and nothing enforces, while a property the author left optional comes back mandatory. Structure — type, enum, required, the object shape — survives the transform, which is why none of it is warned about. One warning per node carries whichever of the three apply, and each list names at most three before counting the rest, because this is a sentence and not an inventory:

Node 'agent' declares a response schema the runtime rewrites before it becomes a grammar: it drops 'maxLength'
rather than enforcing it, requires every declared property ('summary', 'notes' are optional here) and forbids
additional properties.

Since S7 this warning is narrowed at the service, not at the parser. The rewrite it describes is the OpenAI adapter's, and the llama.cpp lane no longer suffers it (§4.2): a node whose model llama-server serves receives the schema as authored. The parser cannot tell which nodes those are — it has no route to the model-to-provider map — so it keeps raising the warning for every Agent node and GraphWorkflowDefinitionService.ValidateAsync drops the ones whose pinned model resolves through ILocalModelProviderResolver to llamacpp. The filter is per warning rather than per node: a node that also earns the pause-context warning above keeps that one. A node with no model pin keeps its schema warning too — what it inherits is decided at run start, and a warning nobody needed is the cheaper error. Nothing else moved: the parser's rule set, the warning sentence and the DTO are all unchanged, and a warning still never blocks.

It warns rather than refuses because a schema is still useful with the constraints in it, the transform is the adapter's business and could change, and every one of these graphs runs. The point is that the author stops believing the parts that do not hold.

The schema is walked breadth-first over exactly the members the transform itself descends — properties, additionalProperties, items, anyOf, oneOf, allOf — so a constraint under items is found and one parked in a $defs or definitions pool is correctly ignored, since the transform never reaches it either. A Start node's inputSchema never warns: nothing compiles it into a grammar, on any runtime. Like every warning it is non-blocking, and the editor's validation strip renders it through the same generic channel as the first kind, so neither the DTO nor the SPA needed a new shape for it. What it is warning about is §4.2.

The node cap runs first of all, ahead of every rule above. MaxNodesPerDefinition reaches the parser as an argument rather than a dependency (GraphWorkflowGraph.Parse(graphJson, maxNodes), defaulted to no cap so the parser stays testable without a container; GraphWorkflowGraphContract.ValidateAndCountNodes passes the option through), and it is checked against the declared length of the nodes array before a single node is read or an edge walked. A cap applied after the parse would bound nothing about the parse that produced it: only the 1 MiB body limit stood between a request and a chain of thousands of minimal nodes, and the acyclicity walk used to spend one stack frame per node, so a deep enough chain overflowed the thread-pool thread's stack — a process kill, not a 400. That walk is now iterative over an explicit stack, which is also what keeps the deliberately uncapped re-parses of a stored graph safe (§3.4). The refusal for a cycle names every node on it in walk order (a -> b -> c -> a), starting and ending at the node the walk came back to.

2.5 A complete example

The eight-node graph below is the module's canonical shape — Start → Agent → Condition → { Pause | Tool } → Parallel → Join → End. It is the React tests' eightNodeGraph fixture (features/graphWorkflows/test/GraphWorkflowFixtures.ts) with every node's position and analyze's timeoutSeconds: null left out for reading — the graph is otherwise the same one.

{
  "schemaVersion": 1,
  "nodes": [
    { "key": "start",  "kind": "Start",  "label": "Start",
      "config": { "inputSchema": null, "defaultInput": null } },
    { "key": "analyze", "kind": "Agent", "label": "Analyze", "maxAttempts": 3,
      "config": { "agentDefinitionId": null,
                  "instructions": "Summarise the request and say whether it needs a human review.",
                  "model": null, "reasoningEffort": null,
                  "responseJsonSchema": { "type": "object",
                                          "properties": { "requiresReview": { "type": "boolean" } } },
                  "includeUpstreamOutputs": true } },
    { "key": "check",  "kind": "Condition", "label": "Needs review?",
      "config": { "path": "output.json.requiresReview" } },
    { "key": "review", "kind": "Pause", "label": "Human review",
      "config": { "prompt": "Approve the analysis?",
                  "allowedDecisions": ["Approve", "Reject"], "requireComment": false } },
    { "key": "lookup", "kind": "Tool", "label": "Read file", "maxAttempts": 3,
      "config": { "toolName": "read_file", "arguments": { "path": "notes.md" },
                  "argumentBindings": { "path": "output.json.path" } } },
    { "key": "fanout", "kind": "Parallel", "label": "Both", "joinPolicy": "Any", "config": {} },
    { "key": "merge",  "kind": "Join",     "label": "Merge", "joinPolicy": "All", "config": {} },
    { "key": "done",   "kind": "End",      "label": "Done",  "joinPolicy": "Any",
      "config": { "outcome": "completed", "resultPath": null } }
  ],
  "edges": [
    { "key": "e1", "from": "start",   "to": "analyze" },
    { "key": "e2", "from": "analyze", "to": "check" },
    { "key": "e3", "from": "check",   "to": "review", "label": "yes",
      "sourceHandle": "true",  "condition": { "op": "Eq", "value": true } },
    { "key": "e4", "from": "check",   "to": "lookup", "label": "no",
      "sourceHandle": "false", "condition": { "op": "Ne", "value": true } },
    { "key": "e5", "from": "review",  "to": "fanout", "label": "approved",
      "sourceHandle": "Approve",
      "condition": { "path": "output.decision", "op": "Eq", "value": "Approve" } },
    { "key": "e6", "from": "lookup",  "to": "fanout" },
    { "key": "e7", "from": "fanout",  "to": "merge" },
    { "key": "e8", "from": "merge",   "to": "done" },
    { "key": "e9", "from": "review",  "to": "done",   "label": "rejected",
      "sourceHandle": "Reject",
      "condition": { "path": "output.decision", "op": "Eq", "value": "Reject" } }
  ]
}

Two details in it are the rules of §2.4 doing their job, and both are easy to get wrong:

  • e9 exists because review offers Reject. A Pause node whose allowed decision has no out-edge that fires on it is refused at save — answering it would strand the run.
  • fanout and done declare joinPolicy: "Any". The check node's two branches are mutually exclusive, so under the default All join each would wait forever for a branch that was never taken, and the run could never reach an end.

e3 and e4 carry no path: they inherit check's config.path.


3. Run lifecycle

3.1 Starting

POST graph-workflows/definitions/{definitionId}/runs answers 202 with the run id. IGraphWorkflowRunService.StartAsync validates, commits a durable intent and signals the dispatcher, which advances the run out of band — so the run legitimately reads Pending when the answer lands.

The caller's requestId is the idempotency key. The lookup before the insert is a fast path two concurrent callers can both pass; the unique index on request_id is the real guarantee, and a reused id naming a different definition is refused rather than answered with a run the caller never asked for.

At start the run pins its own copy of the graph (graph_json and graph_hash on the run row), and StartRunAsync commits the run row, one Pending node run per graph node and the run.created event in one transaction, re-reading the definition inside it so a delete racing the start cannot leave an orphan run. Every node of the pinned graph therefore has a row from the moment the run begins — a Pending row means "not yet judged", not "not yet reached". Editing or deleting the definition afterwards never rewrites a run.

Two checks happen again at start rather than being trusted from save time, because the world moves in between: the graph is re-parsed with the same parser, and every Tool node's tool is re-checked against the live catalog (§4.3). The run also refuses a graph over MaxNodeRunsPerRun and an input over MaxRunInputBytes.

3.2 Run statuses

Pending → Running → WaitingForApproval → Completed | Failed | Cancelled, plus Cancelling (GraphWorkflowRunStatus). There is no Interrupted: runs auto-resume after a host restart and only node runs reconcile. Cancelling exists because cancel is fire-and-forget — the endpoint commits an intent and returns 202, and live node runs drain first.

The run's status is recomputed from scratch at the end of every tick, never accumulated (GraphWorkflowStateMachine.Recompute). It is denormalized so a reader can answer "what is this run doing" without a join. The recomputation is graph-aware, and the ordering matters:

  1. A failed node run outranks everything: the run is Failed, carrying that node's own failure class and reason. When several failed, the one with the lowest node key ordinally is picked, so two readers of the same run cannot disagree.
  2. Otherwise, if a terminal node of the graph (one no edge leaves) Succeeded, the run is Completed.
  3. Otherwise the run is Cancelled, with a reason naming the ends it did not reach and what became of them. A run whose tail was skipped, or whose rejection routed down a branch that skipped the remainder, therefore reads Cancelled rather than Completed like a run that did its job.

While node runs are still live, WaitingForApproval outranks Pending: every node run exists from the start, so there are almost always Pending rows, and reading those as Running would report a run blocked on an unanswered pause as busy — the one thing the two statuses exist to tell apart.

GraphWorkflowFailureClass records why: None — the default, carried by everything that has not failed — then NodeFailed, Timeout, AttemptsExhausted, OutputTooLarge, GateRejected, ValidationFailed, Cancelled, Interrupted, CapacityRejected. GateRejected narrows where a reader should look; it is not a causal proof, because a rejection can route into a branch that then runs perfectly well.

3.3 Node-run statuses and admission

Pending → Queued → Running → Succeeded | Failed | Skipped | Cancelled, plus WaitingForApproval (GraphWorkflowNodeRunStatus). Queued and Running are separate because an admitted node run is not yet executing. There is no Blocked: v1 has no retry routing, so there is no retries-exhausted intervention state to park a row in.

GraphWorkflowStateMachine.Admission answers Wait, Eligible or Skip for a Pending row from its inbound edge states alone. An inbound edge is Satisfied when its source succeeded and its condition fired, Dead when the source settled any other way or its condition did not fire, and Pending until the source is terminal — a source with no row yet is a wait, not a refusal.

A skipped row records one cause, and prefers a branch that broke or was skipped over one a condition merely routed past: a Condition node taking its other branch is the graph working, not news (GraphWorkflowStateMachine.SkipReason).

Retry is in place: Failed → Pending is the one edge out of a terminal node-run status. A failed row under both the node's maxAttempts and the run's MaxTotalAttempts goes back to Pending with the attempt incremented, in one atomic write, and the node.retried event carries the failure the row cleared. Only NodeFailed, Timeout and Interrupted are retryable (GraphWorkflowFailures.IsRetryable) — an over-cap document, a refused gate, a cancelled run, a graph that no longer declares a node and a capacity refusal (CapacityRejected: only the operator ejecting a model or picking a loaded one changes it) all produce the byte-identical answer next time.

A consequence worth stating plainly: a node declaring maxAttempts: 1 reports AttemptsExhausted on its only attempt, because the node's budget is genuinely why nothing will try again. What actually went wrong survives on the row's reason and on its node.failed event.

3.4 The tick

GraphWorkflowDispatcher is one loop with two pumps — a signal channel (bounded, drop-on-full, because a signal is a latency hint) and a sweep every DispatchIntervalMilliseconds over every live run. A dropped signal costs at most one interval of latency, never correctness. The first sweep runs immediately at startup rather than after an interval, so recovery does not pay for one.

Every dispatch-side status write happens inside a serialized AdvanceOnceAsync call. A lane's work produces a pollable result and never transitions a row itself; the only other writers to a run are the human command paths. One tick, in order, and the order is load-bearing:

  1. Steer (§3.7). Judge every unjudged steer entry: a row still Queued/Running on the steered attempt and invocation goes back to Pending on the same attempt; any other entry is recorded ignored. Before the poll, so the poll's superseded-entry sweep cancels the old turn in the same tick.
  2. Poll. Ask every lane what became of the work it was driving and settle what landed — before anything reads the rows for a decision, or the run judges its graph against a row that is only still Running because nothing asked. A row its lane had nothing to say about is offered to its deadline instead.
  3. Drain (only while Cancelling). A drain admits nothing. Every terminal is reached through it or through the "nothing is live any more" recomputation, because writing a terminal over live node runs would strand them under a run no tick looks at again.
  4. Retry the failed rows that still have budget.
  5. Admit the eligible Pending rows, and skip the ones every path into which is dead.
  6. Publish (chat-bound runs only, §3.6): every Succeeded row whose node has publishToChat and no published_message_id yet becomes one chat message. After admit, which settles the inline kinds (End among them), and before the recompute that may end the run.
  7. Recompute the run's own status, against the version it was read at.

The concurrency cap counts working runs, not parked ones. Admission of a Pending run asks CountActiveRunsAsync whether MaxConcurrentRuns are already executing. A run whose live rows wait on a person — at least one WaitingForApproval row, none Queued or Running — is parked and holds no slot, so four chats waiting on their users cannot freeze every new start at Pending behind a 202. Cancelling always counts. A parked run that resumes does not re-pass the admission gate (accepted: the lanes still bound real work). This applies to a Standard run parked on a Pause too.

Deadlines are re-derived from the row every tick — started_at_utc plus the node's timeoutSeconds or DefaultNodeTimeoutSeconds — never armed in memory, so they survive the restart that would otherwise leave a node run bounded by nothing. GraphWorkflowDeadline.Grace adds 30 s before the run ends a row itself: the lane bounds its own turn by the same number from a slightly later moment, and ending the row the instant the number is reached would race a better answer and sometimes win by milliseconds.

3.5 Restart

GraphWorkflowStartupReconciler is an IHostedService registered before the dispatcher, so its pumps cannot admit a row recovery has not judged. It reads the interrupted set — exactly Queued ∪ Running, which the store scopes and the reconciler never widens. WaitingForApproval is deliberately outside it: it is a durable human wait rather than in-flight work, and a reconciler that took it would destroy every pause and chat input on the node on every boot. The reconciler hands its verdicts to the store to apply in one transaction. Exactly-once survives a crash during recovery: a host that dies before that commit leaves the rows as it found them, and the next boot judges them from the same evidence. Recovery makes at most three passes, and the last one settles whatever is left rather than walking away from it.

The verdicts:

  • A Queued row, or a Running inline row, collapses to Pending without touching Attempt. Neither is a failure, so neither costs an attempt.
  • A Running LlmCall, Agent, DecisionModel or Tool row is failed Interrupted. The work was an in-process task with no durable handle, so its partial output died with the host. Recovery never re-attempts it; the dispatcher's retry stage does on its first tick, if and only if the node and run budgets allow. The class written is the plain Interrupted rather than anything GraphWorkflowFailures.Classify would decide, because recovery deliberately never parses the run's graph and so cannot see the node's attempt cap — the retry stage, which does, classifies on that first tick.

That is the same "never resume a provider stream, start a fresh attempt instead" posture Development Mode records in ADR 0001, applied without that design's replacement-attempt rows: a graph workflow node run retries in place with the attempt incremented, because per-attempt history lives in the event log rather than in a second row.

The reconciler touches no run row. The dispatcher's first sweep — PumpSweepAsync runs one immediately rather than waiting out an interval — recomputes the run's status from the rows recovery left behind.

3.6 Chat binding and publishing

A run started from the Chat page (POST graph-workflows/conversations/{conversationId}/messages, Chat "Chat workflow mode") is bound to that conversation: graph_workflow_runs.conversation_id is a real foreign key to conversations with ON DELETE SET NULL, and trigger_message_id names the user message that started it (migration AddChatWorkflowBinding). A partial unique index on conversation_id over the four live statuses (ux_graph_workflow_runs_live_conversation) makes one live run per conversation a database rule; the store answers the losing insert with GraphWorkflowRunBusyException. IGraphWorkflowRunService.StartAsync has an overload taking a GraphWorkflowRunBinding; an unbound run (the Graph Workflows page) publishes nothing. While the newest bound run is live (parked included), a normal chat send into the conversation is refused server-side with conflict type GraphWorkflowRunLiveInConversation, not only by the Chat page's composer lock (Chat "Chat workflow mode"). Every Agent of a Chat graph sees the chat message in its prompt (§4.2, "Conversation request"). A conversation purge (ConversationFootprintPurge) unbinds the conversation's runs explicitly and never deletes one.

Publishing is an outbox pass in the tick (§3.4 step 5), GraphWorkflowChatPublisher: one query for the candidates (ListUnpublishedNodeRunsAsync: Succeeded and published_message_id IS NULL), the node's publishToChat read off the pinned graph, an insert-if-absent of the assistant message under the deterministic id GraphWorkflowChatIds.PublishedMessage(runId, nodeKey, attempt), then MarkNodeRunPublishedAsync — a compare-and-set that stamps graph_workflow_node_runs.published_message_id and appends node.published (detail { messageId }) in one transaction. The stamp does not bump the run's version — it moves no run state, so no run-level write should lose a race to it. Either crash order ends as one message: an insert that landed before a crash is found by its id, and a stamp that landed is never repeated. The message is written with a thin run envelope in the same transaction, so the restart reconcile never backfills it as a chat run. A publish failure is logged and never fails the node; while one fails, the tick skips the recompute (which could be the one that ends the run, after which nothing ticks it again) and counts as work so it re-signals, at most three ticks per run (GraphWorkflowDispatcher.MaxPublishRetries, in memory) — after that the run completes without the message. A run that is cancelled drains without publishing.

3.7 Steer

Operator ruling D4 (2026-09-23): the running Agent or LLM Call node of a chat-bound run can be steered. POST graph-workflows/runs/{runId}/nodes/{nodeKey}/steer with { operationId, message } goes through IGraphWorkflowChatService.SteerAsync, which writes the text as a user message of the run's conversation under the deterministic id GraphWorkflowChatIds.SteerMessage(operationId, runId, nodeKey) and then calls IGraphWorkflowRunService.SteerAsync. An unknown node key is a 404 before anything is written. When the run service refuses, not-found included, the message is removed again. The service commits intent only. It appends { operationId, message, atUtc, attempt, invocationId, steeredBySubject, applied: null } to the encrypted graph_workflow_node_runs.steering_json (migration AddGraphWorkflowSteering, AAD purpose graph_workflow_node_run_steering_json), signals, and answers 202 with the run detail. The append is a compare-and-set that re-checks the run, the row and the cap inside its transaction. It does not bump the run version.

Refusals: 400 for a node that is not Agent or LlmCall, for a run with no conversation, and for a blank or oversized message (the chat-input answer cap, min(MaxMessageSizeKb × 1024, MaxOutputJsonBytes / 2)). 409 GraphWorkflowSteerLimitReached for a row already holding MaxSteersPerNode entries. 409 GraphWorkflowRunConflict for a Pending, Cancelling or terminal run, for a row that is not Queued/Running, and for a reused operationId with a different message or person. The same operationId with the same message is a replay and answers 202 again. Idempotency is per entry, not per attempt, because a steer never changes the attempt.

The dispatcher applies it (§3.4 step 0), because the store's node move has no from-status predicate and only inside the advance gate is "still running" still true when the write lands. A row still Queued/Running on the entry's attempt, and on its invocationId when the steer saw one, moves to Pending with node.steered, and its entries are stamped applied: true in the same transaction. Every applied entry gets its own node.steered event (two steers between ticks are one reset and two events). Any other unjudged entry is stamped applied: false with node.steer-ignored. Both details are { operationId, message, attempt }, so the activity feed renders a running node's steer from the event alone. A Cancelling run ignores them all, and so does a terminal one: a steer can commit after the apply pass of the tick that ends the run, so the terminal early return in AdvanceCoreAsync judges it on the tick the steer's own signal brings. The poll's ForgetSupersededAsync then discards the old flight, and the invocation executor's discard hook cancels its invocation. The next admit re-runs the node on the same attempt: no attempt is spent, MaxTotalAttempts is not charged, and the restart reconciler is unchanged, since a Pending row is not interrupted work. The re-run's seed prompt ends with an ## Operator steering section listing the applied entries, oldest first (§4.2, §4.2a). Each reset restarts the node deadline, so MaxSteersPerNode is the bound. Steering text is not counted against MaxRunInputBytes: each entry is bounded by its own per-entry byte cap (above) and a row by MaxSteersPerNode. Cross-node routing stays absent (register 22, D6(a)).

3.8 Attachments

A chat send to a graph whose chat.acceptsAttachments is on carries references only: run.input.attachments is [{ fileId, name, kind: "text"|"image", bytes }] (bytes is the upload's size), checked with the rest of the run input against MaxRunInputBytes before the user message is persisted. Any other graph answers 409 GraphWorkflowAttachmentsNotAccepted (Chat "Chat workflow mode").

Content is read at attempt time, and only by an Agent or LLM Call node with includeAttachments: true on a chat-bound run (run.input.conversationId). GraphWorkflowInvocationExecutor resolves each reference through IConversationUploadedFileStore:

  • Text — the cached extracted Markdown, composed by the chat's own ConversationAttachmentContextComposer into a ## Attachments section after the prompt (and before any steering section). The section opens with the chat's notice that the content is untrusted DATA, not instructions. Each file is wrapped by UntrustedContentFraming.WrapDocument, with its name inside the fence as file: metadata. The marker's nonce is keyed on the server-secret per-conversation seed (IUntrustedContentFenceSeedProvider) plus the run id and node key, so a document cannot forge the closing marker. Bodies are budgeted before wrapping at MaxRunInputBytes, counted in characters (the composer's unit). Truncation therefore never cuts a closing marker, and the composer's truncation notice follows. A file whose status is not Extracted, or whose Markdown is empty, is skipped, below. The Markdown is read only for an Extracted row, because a .md upload's raw bytes share the Markdown blob's path.
  • Image — attached as a ConversationImagePart on the seed turn through IChatTurnContextBuilder.BuildImageContextAsync, so the chat's MaxImageAttachments / MaxImageAttachmentBytes caps apply (over-cap images are dropped with a warning). Only when the node's effective model resolves SupportsVision; otherwise the node fails ValidationFailed with a reason naming the node, the file and the model. An LLM Call carries images the same way. An image reference whose file's extraction status is not Image is skipped before the vision check, so it never reaches BuildImageContextAsync, which would drop it silently.
  • Skipped — a file no longer in the conversation when the attempt starts, and the unusable text or image files above, are skipped. The same attempt-time resolution that reads the content decides this, so the notice can never disagree with the prompt. The node's output document gains attachmentsSkipped: [name…] (omitted when nothing was skipped), the executor logs the count at Information, and the turn still runs.

A node without the flag, a Standard graph and a run without attachments send a prompt byte-identical to one before attachments existed; every Agent of a Chat graph still sees the Attachments: <names> line of its conversation request (§4.2).


4. The node kinds

Every kind's output goes through one writer, GraphWorkflowDocuments, which composes the common envelope, derives the branch and enforces the size cap. No executor composes a document itself, because a second implementation of any of those is a way for an executor to disagree with the routing the dispatcher will do a moment later.

The output envelope — what an edge condition reads:

{ "status": "succeeded" | "failed" | "skipped",
  "attempt": 1,
  "branch": "approved" | null,   // the label of the first CONDITIONAL out-edge that fires; null when none did
  "output": { /* per-kind, below */ } }

The input document — what an executor is handed:

{ "run":      { "input": /* the run's start payload */ },
  "upstream": { "<nodeKey>": /* that node's output envelope */ },
  "input":    /* the single satisfied predecessor's whole envelope, the upstream map when there
                 is more than one, or null when there is none — which is what Start sees */ }

branch is written even when null, so a reader can tell "no branch fired" from "this document predates branches". An unconditional out-edge never names a branch: it accepts everything, so it says nothing about which way the run went. The composed document is capped at MaxOutputJsonBytes in UTF-8 bytes — a character count would let astral text through at four times the cap — and a node whose document exceeds it fails OutputTooLarge, which is not retryable.

Pause and ChatInput park on a person (GraphWorkflowPauseExecutor); Agent, LlmCall and DecisionModel share the model-invocation lane (GraphWorkflowInvocationExecutor); Tool has its own lane. Start, End, Condition, Parallel and Join are inline (GraphWorkflowInlineExecutor): their work is a pure function of rows the tick has already read, so they run inside the tick with no Queued hop. They still write two rows each, which is what makes the timing of a fan-out visible in the event log.

4.1 Start

Config: inputSchema, defaultInput — both optional raw JSON, and both authoring metadata in v1 (the run's input is validated for size, not against the schema).

Output: { "input": <the run's start payload> }, handed to everything downstream.

4.2 Agent

One headless saved-agent turn, on the model-invocation lane shared with LLM Call and sized at MaxConcurrentRuns (GraphWorkflowInvocationExecutor). It drives IInvocationRunner from the tick and never inside it; the turn is an in-process task with no durable handle, which is exactly why an interrupted Running row is failed rather than resumed (§3.5). InvocationId is written on the row as the correlation id in the node logs for a turn nothing else survives.

The turn's contents come from RunSavedAgentHandler, the node's other unattended caller of this stack, and five of its rules carry over: the locality gate runs before capacity, the capacity reservation is disposed on every terminal path, approval-required tools are stripped from the offer (and the web tools by name, §4.3), IsUnattended is set, and the terminal state is read off a StrongBox<T> the state-changed handler fills rather than off a return value the runner does not have. The executor is a singleton — the lane and its slot count are the node's and outlive both a tick and a DI scope — so every scoped collaborator is resolved inside the task body from its own scope: the scope the tick handed the store is long gone by the time the turn lands.

Config member Meaning
agentDefinitionId Optional saved agent to bind. Null runs an unbound persona from instructions alone.
instructions Required. The seed user turn.
model Optional model pin. Travels as written and is matched against the catalog at run start, exactly as an agent definition's own pin is, so a graph does not become unsaveable because a model was uninstalled after it was authored.
reasoningEffort Optional, one of none, low, medium, high. Checked at save — its vocabulary is closed and cannot go stale.
responseJsonSchema Optional. Must be an object schema ("type": "object"), because the parsed answer lands at output.json and a condition reads a property off it.
includeUpstreamOutputs Defaults true. Inlines the upstream map into the prompt, budgeted at MaxRunInputBytes and truncated with an explicit marker rather than silently.

In a Chat graph every Agent's seed prompt also carries the chat request, whatever includeUpstreamOutputs says: after the instructions and before the upstream map, a ## Conversation request section with run.input.message (same MaxRunInputBytes budget and truncation marker) and, when the run input lists attachments, one Attachments: a.pdf, b.png line (names only; the content reaches an includeAttachments node, §3.8). Agent nodes have no inputBindings, and after a router the upstream is the router's choice rather than the message, so without it an agent downstream of a DecisionModel or Condition never sees what the user asked (live-round finding, 2026-09-23). A Standard graph's prompt is unchanged byte for byte.

A steered row's prompt ends with an ## Operator steering section (§3.7): the row's applied steers, numbered, oldest first, after everything above, the ## Attachments section (§3.8) included. An LLM Call gets the same section after its bound prompt. A row nobody steered sends a prompt byte-identical to one before steering existed.

Output: { "text": …, "json": … | null, "usage": { inputTokens, outputTokens, totalTokens, reasoningTokens, durationMs, finishReason, model } }. Every usage member is nullable: the runner reports what its provider gave it, and no provider reports all of them.

What a response schema actually enforces. llama-server compiles it into a real GBNF grammar and applies that grammar from the first output token, so the structure is guaranteed: the object shape, the declared types, an enum's member list, the required keys. The compiled root rule leaves the model's <think> block optional and unconstrained and forces the schema only on what follows, so a reasoning model is not fighting the grammar.

Since S7 the value bounds are guaranteed too, on llama-server. minLength, maxLength, pattern, format, the numeric minimum and maximum, the item counts and the rest now reach llama.cpp as written, and so does the author's own required list — a property left optional stays optional, and an object that omits additionalProperties stays open. DeferredLlamaServerChatClient.ApplyResponseSchemaPassthrough writes the authored schema onto the request body itself, and the MEAI OpenAI adapter fills its own response format in only with ??=, so its unconditional strict-schema transform is never invoked on this lane. The one exception is a repetition bound above LlamaGrammarToolSchemaCompatibility.MaxGrammarRepetitionBound (1024), which is still stripped because llama.cpp's GBNF converter cannot compile it into a grammar at all — see §8.

On every other runtime the old rule stands. The strict-schema transform relocates twenty-two value keywords into the schema's description string on the way out of the .NET process, where they are advice to the model rather than a constraint (a maxLength: 3 produced a 1302-character field in the S6 live round), marks every declared property required, and closes an object that has properties and omits additionalProperties. There, validate a value bound downstream — in an edge condition or the consuming node — and never in the schema alone. You do not have to notice this unaided: saving or validating a graph raises the non-blocking warning of §2.4 on an Agent node whose schema asks for it, and that warning is now raised only for nodes that are not pinned to a llama-server model. See docs/agent-knowledge.md §2 for the evidence and the --verbose recipe that shows the compiled grammar.

A node declaring a response schema must come back a JSON object. A parse failure fails NodeFailed — the retryable class, since a re-ask under the same grammar can land where one attempt did not — naming the finish reason, because a truncated answer is still Completed and length is the common cause. There is deliberately no salvage path: grammar-constrained output carries no fences, and stripping some would quietly mask a broken grammar. See §8 for what a response schema costs at the llama.cpp grammar layer.

An Agent node is node-local only. If the node's effective model resolves to a cloud model the node run is refused ValidationFailed before any capacity is reserved: an unattended run must never hand the operator's content to a cloud provider. Approval-required tools are likewise stripped from the offer, and the turn runs with IsUnattended set — the same posture a scheduled saved-agent run holds.

The runner's own watchdog reports a timeout as a failed terminal, which is why the executor classes it Timeout rather than NodeFailed — a node that ran out of time deserves a different answer from one whose provider said no. Provider errors never reach the row: the reason says "see the node logs" and the detail stays there.

4.2a LlmCall — LLM Call

One logical chat-model invocation per node attempt, through GraphWorkflowInvocationExecutor.RunLlmTurnAsync and the shared ILocalChatRuntimePackageBuilder / IInvocationRunner stack. The node does not resolve an agent definition or add a persona, tools, skills, memory, or orchestration. Existing bounded transport retries before the first output remain enabled; graph retries are separate attempts governed by maxAttempts. Tool-free requests bypass the function-invocation loop, so unsolicited tool-call content cannot trigger a recovery round.

The package's explicit OmitSystemPrompt flag carries absence through validation, configuration hashing and InvocationAgentFactory.BuildSeedMessages. It defaults to false, so existing chat, Agent and benchmark callers retain their non-empty system-prompt requirement and their existing configuration hashes.

The LLM Call package also requires node-managed llama routing for the selected model throughout dispatch. The shared runtime bypasses cloud selection for that request and rejects a provider remap before choosing a local client, including a cached client. This keeps a settings change between validation and inference from sending the call externally, without locking provider settings for the duration of generation.

Config member Meaning
model Optional installed node-managed GGUF chat model. When omitted, the local-default resolver chooses one. Cloud, external, Ollama, uninstalled and non-chat models are refused before inference. The chosen model is not automatically swapped.
systemPrompt Optional authored system instructions. Empty means no system-role message; no default assistant prompt is substituted.
prompt Required user instructions for this call.
inputBindings Optional object mapping names to dot paths in the node's input document. Values form a named JSON data block alongside the prompt; there is no placeholder expansion or implicit upstream context.
reasoningEffort Optional, using the same vocabulary as the Agent node. Actual support depends on the selected model.
responseJsonSchema Optional object schema, with the same response handling and provider limitations as Agent output.
samplingOptions Optional inference overrides. Unset values use runtime defaults. The editor keeps these in a collapsed Advanced section.

Bindings use GraphWorkflowDocuments.Resolve. A missing path fails the node before inference; an explicit JSON null is a valid bound value. run.input.question reads the run input, and upstream.retrieve.output.result reads an incoming Tool result. input.output.result is a shortcut only when exactly one incoming predecessor is satisfied. The upstream map contains satisfied immediate predecessors, not every ancestor: add a direct dependency edge to consume an earlier result, taking the consumer's join policy into account. Bound data is size-limited and is not silently truncated into a different JSON value.

Sampling overrides include temperature, topP, topK, minP, maxOutputTokens, seed, repeatPenalty, repeatLastN, presencePenalty, frequencyPenalty, stop, and numCtx. The last is a request/prompt budget ceiling; it cannot resize the llama-server context window fixed when the model was loaded.

Output uses the Agent-compatible { text, json, usage } payload inside the normal node envelope. A requested structured response must parse as a JSON object; otherwise the attempt fails. This is not a separate full JSON Schema validation pass. A following Condition can route on output.json.needsReview without parsing prose.

For example, Start → Tool → LLM Call → End can bind document to input.output.result and use the prompt "Summarize the bound document and retain the technical findings." Only the named document is supplied to the model. A second LLM Call can consume the first one's input.output.text the same way.

4.3 Tool

One built-in tool call, on its own lane, through IToolInvocationService (GraphWorkflowToolExecutor).

Config member Meaning
toolName Required.
arguments Optional object of literal arguments.
argumentBindings Optional map of argument name → dot path, resolved against the node's input document. A binding wins over a literal of the same name.
allowedUrls web_fetch only, and required there: the link allow-list (below). Any other tool carrying it is a save error.

Output: { "result": … } — embedded as JSON when the tool answered with an object or an array, and as a string otherwise. That one try-parse is what lets a downstream Condition, which passes its predecessor's output through verbatim, dot-path into a structured answer. Stated rather than hidden: plain text that happens to be a JSON object is embedded as JSON, and the tools that do that mean it.

A Tool node passes two gates, not one (GraphWorkflowToolGate, ruling D6 and ADR 0006):

  1. the tool must be a built-in tool in the ToolCategory.ReadLocal envelope, and
  2. the composed approval policy must answer that it needs no approval.

Both are asked of the same catalog at definition save and again at run start, and the run-start answer wins. A tool that was read-only when the graph was saved does not execute after a policy tightened. A tool outside the envelope is an error, never a warning: a workflow node runs unattended, so a write, execute or approval-gated tool has nobody to ask. Errors are keyed by node key, so the editor draws them on the offending node.

The one web exception (ADR 0017 decision 6; ADR 0006 is its basis and is not amended). While web access is on, a Tool node may name web_fetch if it carries a non-empty allowedUrls list of absolute http(s) URL prefixes (same scheme, host and port; a path prefix on a segment boundary). The list, written before the run, is the consent an unattended node cannot ask for, so there is no pause and no result review. The gate checks the list's shape and any literal url argument at save and start; a bound url, usually upstream model output, is checked at dispatch, which is the real gate: GraphWorkflowToolExecutor calls the fetch service itself with the node's list, and the first URL and every redirect hop must match. Private and local addresses stay blocked even when listed. IToolInvocationService still refuses web_fetch (it is Network), so no other caller reaches it this way. A URL outside the list, a blocked address and web access switched off fail the node ValidationFailed; any other refusal about the page (HTTP error, unsupported type, the fetch's own time budget) succeeds carrying the tool's { "error", "message" } answer, as read_file's own refusal does. web_search is never allowed on a Tool node, and no Agent node — bound agent or default persona — is offered either web tool: both are stripped from the offer by name.

The lane itself is a GraphWorkflowInFlightLane, the same registry the Agent lane uses, and the two differ only in what they run: dispatch to a queue, settle on the poll, a stop that answers no on a repeat, forget what a retry superseded. A tool call passes one queue where an agent turn passes two — there is no node-wide invocation lease to wait for after the lane slot, so the row reads Running the moment the lane hands back an entry. The lane is therefore what bounds the fan-out: a tool call has no global bottleneck of its own, and a Parallel node feeding two hundred search_knowledge_base nodes would otherwise fire all of them at once. It is sized on MaxConcurrentRuns rather than on a knob of its own — the same "how much of this node may be busy at once" question — and is worth a second option only once the two need different numbers. The whole invocation envelope lives inside IToolInvocationService, so the executor enforces none of it and cannot skip any of it: every refusal arrives as an outcome and becomes a row.

A landed outcome maps to a terminal: Executed succeeds; UnknownTool, NotInvocable and InvalidArguments fail ValidationFailed and are therefore never re-attempted; Timeout and Faulted fail on the two retryable classes. The service's own reason is repeated verbatim and is structural by contract — it never echoes an argument value. An output document over the cap is a real, reachable outcome (a knowledge-base search may legitimately answer with fifty thousand characters) and is deliberately not retryable, since the same call composes the same bytes; the message names the tool beside the node, because "which node" alone does not say what to shrink.

graph-workflows/tools is the picker's feed, the same list the save gate checks (GraphWorkflowToolGate.ListAsync: the invocation envelope, plus web_fetch flagged requiresAllowedUrls while web access is on), so the picker cannot offer a name the run would then refuse. A missing binding path fails the node ValidationFailed; the Queued write happens before bindings resolve, so a refused row keeps its input document.

4.4 Condition

Config: path — the optional node-level default dot path its own out-edges inherit.

Output: a verbatim pass-through of the predecessor's output. This is what makes a Condition node a real router: edge conditions evaluate against the source node's output document, so without the pass-through a Condition's own out-edges would inspect {} and never fire. A node with several predecessors has no single upstream output to carry forward and answers {} rather than inventing one.

4.5 Parallel and Join

No config — they are shape, not settings, and §2.3 is the whole of their semantics.

Parallel passes its predecessor's output through exactly as Condition does. Join answers a per-source map over its satisfied inbound edges, { "<nodeKey>": <that node's envelope> }, so everything downstream of a join sees every branch rather than whichever one the single-predecessor shortcut would have picked.

4.6 Pause

Config member Meaning
prompt Required. What the person is being asked.
allowedDecisions Required, non-empty, distinct, from Approve / Reject. Empty would be a question nobody could answer. Answer is refused: it belongs to ChatInput (§4.8).
requireComment Defaults false.

GraphWorkflowPauseExecutor is the one lane that drives nothing: parking a row on a human is two status writes, so there is no work to hold, no slot to wait for and no answer to poll for. It writes Running then WaitingForApproval with PendingDecisionKind naming the pending act, and no output document — a pause's output is its answer, and it has none yet. The prompt, the allowed answers and requireComment are not copied onto the row: they are already in the pinned graph, and a second copy is a second thing that can drift.

The answer arrives through POST .../runs/{runId}/nodes/{nodeKey}/decide, idempotent on a client-minted operationId. Both answers succeed the node run — the answer is the node's output, and routing on it is the edges' job. A rejection reaches the run through an out-edge, never through a node failure. WaitingForApproval can never move to Skipped: skipping an open pause would be an operator walking past a decision instead of giving one, which is the one thing a pause exists to make impossible.

Output: { "decision": "Approve" | "Reject", "comment": … | null, "payload": … }. decision is the enum's name, because that is the member every out-edge condition selects on and the exact member the definition-time pre-flight check writes; the two spellings must produce the same string or a graph that pre-flighted clean would route nowhere. comment and payload ride beside it rather than above it: they are the operator's, not the router's, and a condition that could select on them would route on free text.

A Pause's successor gets only the decision, and the editor wires around it. A node's input is its ONE satisfied predecessor's output (§3.2), so an authored X → Pause → Y hands Y the approval and never X's answer, and a Pause before an End loses the result the same way. Two things address that, neither of them a change to what a Pause writes. The validator raises the non-blocking warning of §2.4 on Y. And the editor adds the missing edge itself: when an author connects an edge into or out of a Pause, pauseContextEdges returns the unconditional edges labelled context from the pause's nearest non-Pause ancestor to its successors, the canvas adds them, and a notice says one appeared. It descends from the Open Canvas importer's CanvasWorkflowImport.AddPauseContextEdges (§9.2) — walking back through consecutive pauses, skipping a self-loop — judged per successor exactly as the validator judges it: only a successor whose every inbound edge leaves a Pause is starved (one something else already feeds is not, so it gets nothing and no warning), and the ancestor is the union of the nearest non-Pause ancestors of every pause feeding it. The importer applies the same rule and the same guards (§9.2), so the three halves cannot advise different edges. A Condition ancestor is skipped, because the added edge carries no sourceHandle and would save as that Condition's second unconditional out-edge, which §2.4 refuses. A pause whose nearest non-Pause ancestor is not unique (mutually exclusive branches feeding it) gets nothing, because wiring both ancestors into an All successor would skip it the moment the untaken branch is dead. And a successor with joinPolicy: "Any" — of any kind, since the policy is a property of every node (§2.3) — gets nothing, because an unconditional content edge would admit it while every approval was rejected. Wiring into a pause visits each pause reachable forward through consecutive pauses with its own ancestry, so A → P1 → P2 → B plus X → P2 adds nothing for B (its ancestors through P2 are {A, X}). Everywhere else Y keeps its default All join policy, so it is admitted only once both the content and the approval have arrived, and its input is the upstream map carrying both.

The pass runs on the connect gesture and nowhere else — never on render, never on validate — so an edge the author deletes stays deleted. Wiring out of a Pause considers only the node just connected, since that is the only successor that gesture can have starved; wiring into a Pause considers every successor of it, walking forward through consecutive pauses so A → P1 wired last still reaches the B behind P1 → P2 → B.

A decide call carrying a different operationId for an already-answered pause is 409 GraphWorkflowGateAlreadyDecided, with the standing decision on the body — a second human act is refused, not replayed.

4.7 End

Config member Meaning
outcome Required. The declared outcome string.
resultPath Optional dot path into the End node's input document.
publishToChat Defaults true in a Chat graph, false otherwise; true is refused outside a Chat graph (§2.1).

Output: { "outcome": …, "result": … } — the resolved path, or the whole input document when the author named none. A path the document does not carry resolves to null; failing the node instead would end a run that did all of its work over a projection nobody reads.

An End node is a terminal node by construction (nothing may leave it), and a run is Completed only once one of them succeeded.

4.8 ChatInput

Config member Meaning
prompt Required. What the chat surface shows while the run waits for the user's next message.

Legal only in a Chat graph, and it needs at least one unconditional out-edge (§2.4). It rides the pause lane: GraphWorkflowPauseExecutor owns Pause and ChatInput alike and writes PendingDecisionKind from the node kind — Approve for a pause, Answer for a chat input — so a reader tells a gate from a question off the row, without the graph. The run reads WaitingForApproval (the wording is the surface's, derived from the pending kind); a restart leaves the row alone exactly as it leaves a pause, and the cancel drain cancels it the same way.

The answer arrives through the same decide route with decision: "Answer" and payload: { "text": "…" }. GraphWorkflowStateMachine.IsDecidable(kind, status, decision) is the one place the pairing lives: a ChatInput takes Answer and nothing else (409 for Approve / Reject), and a Pause refuses Answer (409). A missing, blank or non-string payload.text, or a comment beside it, is a 400; so is text over min(Security:MaxMessageSizeKb × 1024, MaxOutputJsonBytes / 2) UTF-8 bytes — the answer is a chat message, and it must still fit the envelope it is embedded in. operationId idempotency and the 409 GraphWorkflowGateAlreadyDecided naming the standing answer work exactly as for a pause.

Output (GraphWorkflowDocuments.ChatInputOutput): { "decision": "Answer", "text": … }. decision is kept because the replay and standing-conflict checks read output.decision structurally. An edge may condition on output.text. The §2.4 pause-context warning and the editor's pauseContextEdges gesture treat a ChatInput like a Pause: a successor reached only through one receives the answer document, not the content before the wait.

4.9 DecisionModel

Config member Meaning
question Required. What is being decided.
labels Required. 2–32 distinct, non-blank strings of at most 64 characters. Flat on purpose: they become a grammar enum, far below the repetition bound.
provider Optional closed vocabulary, llm only (and the default). A provider this build cannot run is refused at save.
model Optional, as on LlmCall: an installed node-managed GGUF chat model, else the local default.
inputBindings Optional, exactly as on LlmCall (§4.2a).

A classifier node on the invocation lane. The named IGraphWorkflowDecisionProvider (Services/GraphWorkflows/Decisions/) lowers the node to an LlmCall config — a fixed classifier system prompt, the question plus the label list as the prompt, the same bindings, reasoningEffort: "none", temperature: 0 (so the same input routes the same way on a re-run) and the response schema { type: object, properties: { choice: { type: string, enum: labels } }, required: [choice] } — which runs through RunLlmTurnAsync unchanged, model gate, capacity and binding rules included. The provider then interprets the answer, and the executor holds the choice to the labels: an answer that is not a JSON object, or names anything outside them, fails NodeFailed (retryable, so AttemptsExhausted on the last attempt), and the reason never repeats the model's text. LlmDecisionProvider is the one provider; Laya/ONNX classifiers are later providers of the same seam.

Output: { "choice": "<label>", "confidence": null, "probabilities": null, "provider": "llm", "usage": { … } }. confidence and probabilities are always null from the llm provider — a grammar-constrained answer carries no calibrated score, and inventing one would route on noise. Out-edges route on output.choice, directly or through a Condition. An interrupted Running DecisionModel row is failed Interrupted at startup like every model turn (§3.5).


5. API and hub

Full route table and hub inventory: API & Hubs. LocalApiRoutes.GraphWorkflows is the family; the whole surface, hub path included, sits behind request-path middleware in Program.cs that answers 404 when GraphWorkflows:Enabled is false — ahead of the security middleware, so the switch cannot be probed by status code. Every route is Operator-gated.

Route Notes
graph-workflows/definitions GET lists without the graph blob (it is the encrypted column), each row carrying its denormalised kind; POST creates.
graph-workflows/definitions/{definitionId} GET / PUT with the version it was edited from / DELETE, which 409s while a live run pins the definition.
graph-workflows/definitions/validate POST a graph and get its errors and warnings back without saving. The editor asks the runtime's own parser. valid is still zero ERRORS: a graph that only warns passes here and saves.
graph-workflows/tools The Tool node picker's feed (§4.3).
graph-workflows/definitions/{definitionId}/runs POST start → 202 with the run id; requestId is the idempotency key.
graph-workflows/runs The run list, newest first, ?status=&limit= with limit required and capped at 200.
graph-workflows/runs/{runId} One run, its node-run summaries, the run's own resolved output, and the graph this run pinned at start. No node-run documents — those are a per-node read.
graph-workflows/runs/{runId}/cancel 202. Live node runs drain first, so the run reads Cancelling. A repeat cancel is an idempotent 202.
graph-workflows/runs/{runId}/nodes/{nodeKey} One node run in full, input and output documents included.
graph-workflows/runs/{runId}/nodes/{nodeKey}/decide Answers a pause (§4.6) or a chat input (§4.8).
graph-workflows/runs/{runId}/nodes/{nodeKey}/steer POST { operationId, message } → 202 with the run detail. Steers the queued or running Agent/LLM Call node of a chat-bound run (§3.7). 400 for the request, 404 for an unknown run or node, 409 GraphWorkflowSteerLimitReached at the cap, 409 GraphWorkflowRunConflict for the run's state or a reused id. The node-run read carries steering: [{ operationId, message, atUtc, attempt, applied }].
graph-workflows/runs/{runId}/events The event log, paged from an exclusive afterSeq, capped at EventReplayLimit — which the response reports rather than leaving a client to infer it from a full page.
graph-workflows/conversations/{conversationId}/messages POST a chat send into workflow mode → 202 { runId, messageId, action } (§3.6, Chat). 409s: GraphWorkflowRunBusy, GraphWorkflowRerunConfirmationRequired, GraphWorkflowAttachmentsNotAccepted.
graph-workflows/conversations/{conversationId}/runs GET the conversation's bound runs newest first (?limit=, default 20): run summary, definitionId, definitionName (null once deleted), triggerMessageId, pendingInput, steerable.

Six routes cap the request body at 1 MiB (GraphWorkflowRequestSizeLimit): create, update and validate, which carry a graph, start-run, which carries an input document, the chat send and the steer. Without it they would inherit Kestrel's 30 MB default and a body that size would be parsed, walked and hashed before the node cap could refuse it. Kestrel enforces the cap as it reads, inside model binding, so the 413 comes from RequestBodyTooLargeExceptionHandler rather than from the endpoint. A name is capped at 200 characters, a description at 1024.

The run detail carries the run's own graph. GraphWorkflowRunResponse is (Run, NodeRuns, Output, Graph). Graph is the copy this run pinned when it started, not the definition's current one, in the same wire shape GraphWorkflowDefinitionResponse.Graph carries — so a client parses a definition and a run with one piece of code. It sits on the run DETAIL and deliberately not on the run summary: a list must not carry one graph per row, and a graph may hold up to a mebibyte. output is the run's resolved result, written once by the succeeded End node at terminalization; the per-node input and output documents remain a separate read. Without the pinned graph a run view drew the definition it names, which is the wrong graph for every run started before an edit — the whole reason a run pins a copy at all.

A pinned graph that will not deserialize throws, where a node-run document reads as null. The difference is what each blob is: a node-run document is written by the runtime and a broken one is worth reading a page about, while a graph was parsed before it was ever stored, so no supported route reaches a corrupt one. An empty canvas drawn beside a real nodeCount would report the corruption as a graph nobody drew.

Validation errors are keyed. GraphWorkflowValidationException carries a GraphWorkflowValidationResult — a list of (key, message) pairs where the key is the node or edge, or null for a failure about the document as a whole — and the endpoints replay them one by one instead of collapsing them into a sentence. It is therefore kept out of DomainValidationExceptionHandler, which maps single-message validation exceptions globally.

Warnings travel on the validate response as a second list of the same (key, message) shape. They are a separate Warnings member on GraphWorkflowValidationResult — never mixed into Errors, which is what the exception path and valid read — so nothing that refuses on the errors has to filter them out, and a client that ignores the member behaves exactly as it did before it existed.

The hub

GraphWorkflowRunHub at /api/local/v1/graph-workflows/hub. SubscribeRun(runId, afterSeq) joins the per-run group graph-workflow-run-{runId:N} before reading the replay, so a change published between the read and the join cannot reach nobody; the overlap that creates is harmless, because every push is idempotent and keyed by sequence. The snapshot carries the run status, the queued / running counts, pendingDecisions (parked rows whose PendingDecisionKind is not Answer — a pause) and pendingInputs (parked rows waiting on an Answer — a chat input), the watermark, up to EventReplayLimit events, and a replayTruncated flag read from one row past the limit rather than inferred from a full page. There is no in-memory buffer — the store is the replay authority — and a disconnect cancels nothing, because a run outlives the tab. UnsubscribeRun(runId) leaves the group, and StreamNodeActivity(runId, nodeKey) streams one running node's live turn (below).

The pushed event is graphWorkflowChanged, carrying (runId, seq, kind) and no content at all: the subscriber re-reads the named feed from its own watermark, so a dropped push degrades to a late read rather than to a wrong render. kind is lowercase on the wire — run, node, gate — written as literals in GraphWorkflowEventPublisher.ToWireKind and asserted literally on both sides. There is no event kind: every kind moves the append-only event feed, so the client invalidates it unconditionally.

The publisher is a store decorator (PublishingGraphWorkflowStore), so exactly one ping is emitted per committed mutation that has a subscriber, carrying the sequence that commit allocated — and every published transition carries a fresh, increasing sequence. Three writes deliberately publish nothing, because nothing is subscribed to them: the definition writes (CreateDefinitionAsync, UpdateDefinitionAsync, DeleteDefinitionAsync), StartRunAsync, and the startup reconciler ReconcileNonTerminalNodeRunsAsync. API & Hubs states the same. A publisher failure is logged and never fails the write that already committed. Client.Application depends only on IGraphWorkflowEventPublisher; the host swaps the hub-backed implementation in over a registered no-op, so a host without the hub stays resolvable.

The event vocabulary is the closed nineteen-token GraphWorkflowEventTypes catalog — run.created, run.started, run.waiting, run.completed, run.failed, run.cancelled, node.queued, node.started, node.completed, node.failed, node.skipped, node.cancelled, node.interrupted, node.retried, gate.requested, gate.decided, node.published (amendment 2026-09-23, Chat Workflows S1: a result became a chat message, §3.6), and node.steered / node.steer-ignored (amendment 2026-09-23, Chat Workflows S3: a steer reset its row, or arrived after the row settled, §3.7). The feed is append-only and durable, so a token written once is a token every later reader must understand: extend it by amendment, never silently. Event details are small structured payloads — a failure summary, a decision outcome — and never a transcript.

Live activity (Chat Workflows S2). StreamNodeActivity(runId, nodeKey) is a server stream over the node run's InvocationId, served straight from IInvocationResumeRegistry.ResumeAsync: a snapshot, then offset deltas for content, reasoning, tool calls and phases, then the terminal event — the same ChatStreamEvents LocalChatHub.ResumeMessage serves, so the client reuses the chat stream folding. The registry already mirrors every invocation from IWorkerEventDispatcher.InvocationStateChanged, graph turns included, so there is no new publisher and no content-bearing group: graphWorkflowChanged stays content-free. The stream is not tracked as a chat attachment (LocalChatHub.TrackAttachment), which would mark the graph turn detached for the reaper. It refuses with a HubException when the run is unknown, the node key is not a row of this run, the row is not Running, it has no InvocationId, or the registry no longer holds the turn; a client re-subscribes when a node ping changes the running row's InvocationId. The registry is in memory, so a mid-node restart loses the live text: the row is failed Interrupted and retried by the existing rule (§3.5), and the node's durable output is what remains.


6. Frontend

Client.React/src/features/graphWorkflows/ — see React Client for where it sits among the feature folders. One route, /graph-workflows, with all four selections (definitionId, runId, nodeKey, tab) as search params rather than path segments, so every view is linkable and a reload lands back on it. The route component is a thin adapter; GraphWorkflowsPage itself is router-free and is rendered directly in unit tests.

Directory Holds
models/ GraphWorkflowModels.ts (the one file naming generated DTOs, plus the closed vocabularies and narrowers), GraphWorkflowCanvasModels.ts (the discriminated node union and the graphToCanvas / canvasToGraph round trip), GraphWorkflowLayout.ts, GraphWorkflowValidation.ts (the client mirror of the graph rules), GraphWorkflowRunGraph.ts.
queries/ Every read and mutation over the generated adapters, including the forward-paged events feed, the several-node-runs read (useSettledGraphWorkflowNodeRuns) and the two chat routes (useGraphWorkflowConversationRuns, useSendGraphWorkflowChatMessage).
hooks/ useGraphWorkflowEditor (controlled React Flow state, per-handle connect prefill, refusal of a second unconditional edge, and the context edge added around a Pause or ChatInput on connect — §4.6) and useGraphWorkflowRunHub.
components/ Editor: the per-kind node cards, the canvas with its palette and Auto-arrange, the validation strip, the node and edge config panels (the DecisionModel body and the input-bindings list shared with LlmCall under config/), the workflow settings popover (graph kind and the chat block, saved with the graph), the definition list and meta dialog. Run view: the status badge, the read-only run graph, the node-run table, the run list, the events tab, the node panel and the decision panel (Approve/Reject buttons for a Pause, a text-answer form for a ChatInput).
samples/ The sample graphs the New workflow dialog can start from, and their index GraphWorkflowSamples.ts (see Samples below).
pages/ GraphWorkflowsPage — editor mode without a runId, run mode with one.
api/ GraphWorkflowConflict.ts, which reads the NodeConflictProblemType members by name — the three run/definition/gate refusals and the three chat-send ones (GraphWorkflowRunBusy, GraphWorkflowRerunConfirmationRequired, GraphWorkflowAttachmentsNotAccepted).
features/chat/workflow/ (the chat side, not this folder) useChatWorkflow (the chat page's workflow mode: picker state, send routing, 409 handling, the live bound run's hub subscription), WorkflowSelectorCard, WorkflowRunStatusCard, WorkflowActivityBlock, ChatWorkflowModels.ts (taken path by layout rank, activity rows), ChatWorkflowStore (UI state only). It imports this feature's queries, hub hook, layout, models, status badge and conflict reader — ten reviewed no-cross-feature fingerprints in config/dependency-baseline.json. See Chat.

The editor mirrors the server's rules rather than inventing its own, and the mirror is a mirror on purpose: the authoritative answer comes from graph-workflows/definitions/validate, which runs the parser a run would run, and serverErrorsToIssues maps its keyed errors back onto the canvas elements. The page flow is Validate → Save (server validate first; a 409 reloads) → Start. The server's warnings arrive through serverWarningsToIssues as issues with the rule serverWarned and reach the strip through its own warnings prop; they render in their own alert below the errors, in a different colour and under a title that asks rather than refuses, and never join the list the save gate, the red chips and the config panels read, because a warning is not a reason to save less. Every client rule mirrors a rule that REFUSES a save, so the client raises none of them.

Two round-trip rules the editor has to hold because the server does. Every JSON-shaped config field is written back as JSON text with a string quoted, since that text is exactly what a save parses again — an unquoted string read as the invalid-JSON complaint on a defaultInput the server had accepted. And the client's dot-path mirror refuses an EMPTY segment the way GraphWorkflowTokens.IsDotPath does, so a..b cannot read green in the drawer and then 400 on save.

Samples. The New workflow dialog offers "Start from" with a blank workflow or one of the samples in features/graphWorkflows/samples/: one wire graph per <id>.json, indexed by GraphWorkflowSamples.ts with an i18n name and description. Picking one prefills the name and description (never over text the operator typed), and the create call posts the sample's graph instead of the Start → End starter, so a created sample is an ordinary definition with no link back to its file. Every sample must run on any installed node-managed GGUF chat model and nothing else: no Tool or ChatInput node, no model pin on any node, a Standard graph, and a Start defaultInput so the Start Run dialog is one click. An Agent node is allowed only without agentDefinitionId: unbound, it runs the built-in default persona with the local tools a fresh install already has (§4.2), which is how agent-research-and-check shows where an agent fits (Agent → LLM Call check → Pause → End); a bound agent would need setup first. Prompts inside the graphs are model-facing config and stay English. GraphWorkflowSamples.test.ts holds the client half (validator, lossless canvas round trip, bundle keys), and XE-Local-AI-Engine.Tests/GraphWorkflows/GraphWorkflowSampleContractTests.cs parses every file through the real GraphWorkflowGraph.Parse and also refuses any warning, so a sample never opens with a note the operator did not cause.

useUnsavedChangesGuard is set with allowSameRoute, so writing a search param — selecting a node or a tab — does not trigger the leave prompt, while a real route change still does.

Nodes are laid out left-to-right by layoutGraphWorkflow, a dependency-free layered DAG layout (longest-path ranking in Kahn order, one barycenter pass, ties broken by node key). It is used for three things: a node that arrives without a position, the editor's Auto-arrange, and the run view's nodes-only fallback for a response that carries no graph at all. A definition whose nodes carry no positions opens laid out and dirty — the layout is a proposal the operator has not saved yet. That is by design, and §9 is why it matters.

The run view is read-only, and it draws the graph the run pinned at start — the one GET runs/{runId} carries (§5) — with each node run's state overlaid on it. A definition edited since the run started is then worth saying and nothing more: an informational notice that this is the graph the run itself ran on, never a reason to draw less. Nodes only, auto-laid-out, is what remains for a response that carries no graph and whose hash disagrees with the definition on screen, because drawing today's edges over an older run would be a lie about routing. The Pause decision panel reads its prompt, its allowed answers and requireComment off the pinned graph for the same reason: a definition edited since could offer a decision this run's gate would refuse. The view also lists the node runs, shows a selected node's input, output and error documents, and renders the event trail. The events feed pages forward from afterSeq: 0, so replayTruncated means the newest events are missing, and the banner says exactly that.

The graph → canvas conversion is memoised on the graph object. graphToCanvas parses the whole document and lays out every node that carries no stored position, while the run hub invalidates the run detail on every event — so one node transition would otherwise re-parse and re-rank up to a mebibyte of graph to redraw a single badge. The cache is a WeakMap keyed on the graph object itself: the conversion is a pure function of that document, TanStack's structural sharing hands back the same object while its JSON is unchanged, and an entry dies with the graph that keyed it. Callers treat the cached canvas as frozen and copy every node and edge they annotate.

useGraphWorkflowRunHub degrades rather than fails. On a subscribe it re-sends its watermark as afterSeq; the store serves a client that has been away for days, so there is no buffer to roll over. While the hub is unavailable it hands the page a poll interval of 3 s, cleared only on a successful subscribe. Node-run invalidation is keyed on the run, not on the selected node, so clicking through the node table does not tear down the subscription.

The feature ships on and is a top-level navigation entry. Its capability flag, nodeCapabilities.graphWorkflows, is declared in src/capabilities/NodeCapabilities.ts with the rest; the route redirects home when it is off.


7. Testing

See Writing Tests for which project a new test belongs in and Testing & Validation for how to run them.

Layer Where
Parser, conditions, documents, state machine, deadline, dispatcher, executors, run service, options XE-Local-AI-Engine.Tests/GraphWorkflows/
Endpoints XE-Local-AI-Engine.Tests/Endpoints/GraphWorkflows/V1/
Hub and event publisher XE-Local-AI-Engine.Tests/Hubs/GraphWorkflow*Tests.cs
Store, encryption, purge coverage XE-Local-AI-Engine.Client.Persistence.Tests/GraphWorkflows/
React components, hooks, models, queries colocated *.test.ts(x) under features/graphWorkflows/
Browser end-to-end XE-Local-AI-Engine.Tests.E2ETests/Tests/GraphWorkflowE2ETests.cs

Two suites are worth naming because they are the ones that catch a wiring break the unit tests cannot:

  • GraphWorkflowEndToEndTests and GraphWorkflowPauseToolEndToEndTests drive the wired host over REST and a real SignalR HubConnection, with the model runtime replaced by Testing.FakeOllama. The pause/tool suite asserts the gate push twice, which is what proves the publisher decorator and the lowercase wire kind agree with the client.
  • The E2E class drives the real browser through the whole loop — create a definition, drop nodes from the palette, wire and configure them, validate, save, start a run, answer the pause, read the event tab.

GraphWorkflowHarness and the *HostFixture types are the shared seams; GraphWorkflowGraphs holds the reusable graphs. On the client, test/GraphWorkflowFixtures.ts is the one fixture file, and eightNodeGraph is the graph of §2.5.


8. Limits and options

GraphWorkflowOptions, bound from the GraphWorkflows configuration section. No appsettings*.json carries the section — the defaults below are the shipped values, and an operator overrides one through configuration or an environment variable such as GraphWorkflows__MaxConcurrentRuns.

Option Default What it bounds
Enabled true The whole surface. False makes every route and the hub answer 404 through the request-path middleware, while registration stays intact so a disabled node answers legibly instead of 500-ing out of an empty container.
MaxNodesPerDefinition 200 Nodes in one definition, enforced when it is validated rather than when it runs, and checked against the declared node count before the graph is parsed (§2.4). A run that re-parses an already-stored graph is deliberately uncapped.
MaxNodeRunsPerRun 200 Node runs one run may instantiate. Never below MaxNodesPerDefinition — a run that could not instantiate the definition it started from would fail halfway through a graph the operator was allowed to save.
MaxTotalAttempts 50 Every attempt one run may spend across all its nodes. The guard against a retry storm.
DefaultNodeTimeoutSeconds 600 One node run's attempt, when its node names no timeoutSeconds. Unlike Dev Workflows, a node that declares nothing still has a deadline.
MaxOutputJsonBytes 262 144 One node run's composed output document, in UTF-8 bytes, checked before it is encrypted and stored.
DispatchIntervalMilliseconds 500 The sweep cadence, independent of the change signals the dispatcher also listens for. Floored at 100 ms.
MaxConcurrentRuns 4 Executing runs at once — a run parked on a person holds no slot (§3.4) — and the size of both the shared Agent/LLM Call/DecisionModel invocation lane and the Tool lane. Runs above the cap wait; they are not refused.
MaxRunInputBytes 65 536 A run-start input document, checked in GraphWorkflowRunService.StartAsync (§3.1) rather than at the endpoint, so every caller of the service is held to it. Also the budget for the inlined upstream map in an Agent prompt and the complete user prompt with bound JSON data in an LLM Call.
EventReplayLimit 200 Events one replay may return, hub snapshot and events route alike. Ceiling 1000 — one replay is one response body.
MaxSteersPerNode 5 Steers one node run may take (§3.7). Each applied steer re-runs the node on the same attempt and restarts its deadline, so the attempt budget does not bound steering; this does. Floored at 1.

GraphWorkflowOptionsValidator checks at startup what the data annotations cannot: a semantic floor under each budget (a MaxNodesPerDefinition of 1 passes [Range(1, …)] and still admits no graph, since every graph carries a Start and an End), the MaxNodeRunsPerRun ≥ MaxNodesPerDefinition relation, and the replay ceiling. An operator meets these at boot rather than once per node run.

Two limits that are not options, because they are not runtime budgets: the 1 MiB request-body cap on the five routes that carry one (create, update, validate, start-run and the chat send), and the 200-row cap on a run list page.

One ceiling that lives outside this module: an Agent node's responseJsonSchema goes down the same llama.cpp GBNF path as a tool schema, which has an empirical combined repetition bound (LlamaGrammarToolSchemaCompatibility.MaxGrammarRepetitionBound, 1024). Keep a response schema flat. Since S7 that sanitiser covers the response schema too: DeferredLlamaServerChatClient.ApplyResponseSchemaPassthrough runs the authored schema through the same Sanitize pass before it reaches the wire, so an over-large bound is dropped rather than failing the turn with HTTP 400 Failed to initialize samplers. Every bound within the cap is now enforced by the grammar rather than dropped — see §4.2. See Local Runtime & Providers and docs/agent-knowledge.md §2.


9. The Open Canvas import

Open Canvas (the Preview workflow builder) was removed. Its saved canvases are converted into Graph Workflow definitions automatically during startup. There is no button, prompt or opt-out. The encrypted source survives interrupted conversion; startup stops on conversion failures and retries after the underlying fault is resolved.

Back up the node's data directory before upgrading a node that has canvases you care about. That is not a formality: see the failure posture below.

9.1 Why it is split around the migration

An EF migration cannot decrypt graph_json — it has no node key, and the column is AEAD ciphertext — so the conversion has to run in application code. But the DropCanvasWorkflows migration removes canvas_workflows during ApplyNodeChatMigrationsAsync, which runs before any hosted service starts. A single post-migration importer would therefore always find the table gone.

So the step is split, in Program.cs:

var pending = await ReadPendingCanvasWorkflowsAsync(app.Services);   // read + decrypt, BEFORE migrations
await ApplyNodeChatMigrationsAsync(app.Services);                    // backs up, then runs DropCanvasWorkflows
await ImportCanvasWorkflowsAsync(app.Services, pending);             // write, IMMEDIATELY after that pass
await ApplyNodeIdentityMigrationsAsync(app.Services);                // a different database; runs after the write

Before migrations, the reader uses one SQLite transaction to copy the legacy rows into canvas_workflow_import_recovery without decrypting or rewriting their graph bytes. This is a durable staging table in the node database, not a SQLite temporary table or a restored Open Canvas feature. While the original table exists it remains authoritative, and a retry atomically refreshes staging from it. Once migrations drop the original, startup reads staging instead. A process interruption between migration and import therefore loses no source rows.

After migrations, the importer uses the same scoped NodeChatDbContext as the definition service and store. It writes all definitions and drops the recovery table in one transaction. A write, cleanup or commit failure rolls back both operations; a later startup imports the retained ciphertext without duplicate definitions. Fresh installs have neither table and import nothing. An empty legacy table is staged and cleaned up without creating definitions.

The import runs regardless of GraphWorkflows:Enabled, deliberately: an operator who never turns the feature on must not silently lose their canvases. Every row is read — there is no cap — because a cap plus an unconditional drop in the same build would destroy everything past it as its normal outcome.

Nothing in the importer may depend on a Preview type: the read parses the stored blob into private records of its own, so the Open Canvas namespace can be deleted out from under it. It also runs under --reset-admin-password, whose branch of Program returns only after the migration pass, so the read, the migrations and the write all happen first. The knowledge-downgrade commands are the one launch path that returns before migrations, and therefore before the import.

9.2 The mapping

Open Canvas node Graph Workflow node
Start Start, with the canvas StartText as config.defaultInput = { "text": <StartText> } (null when empty) — an object, because the editor renders a stored default as JSON text and a bare string would never re-parse.
Agent Agent with maxAttempts: 1, instructions / model / reasoningEffort carried over, agentDefinitionId: null, responseJsonSchema: null, includeUpstreamOutputs: true.
Debug Elided. Every X → Debug and Debug → Y collapses to X → Y. A Debug node was a side-event tap that forwarded its input unchanged, so removing it preserves the run's meaning exactly.
Pause Pause with prompt: "Approve and continue?", allowedDecisions: ["Approve"], requireComment: false; its single out-edge gains label: "approved" and condition: { "path": "output.decision", "op": "Eq", "value": "Approve" }. Open Canvas's Continue was a resume rather than a decision, so Approve alone is the faithful translation — and one allowed decision with one matching out-edge satisfies the Pause pre-flight rule of §2.4. Plus one context edge X → Y, unconditional and labelled context, where X is the pause's unique nearest non-Pause ancestor along the (Debug-elided) chain and Y each starved successor of it — see below.
End End with outcome: "completed", resultPath: null.
ModelProfile (on an Agent) Dropped. An Agent node's config has no profile member. A non-null value records a reason naming the canvas id and the node key, never the value. That reason reaches the log as a per-change Warning only when the definition imported cleanly; on the needs-attention path it is folded into the description instead (§9.3).

Node keys are the Open Canvas ids sanitized to [A-Za-z0-9_-]{1,64}; an unmappable or duplicate key becomes n{index}, and a sourceId → key map rewires the edges. Each emitted edge gets a minted key e{index}. Node and edge keys share one namespace, so CanvasWorkflowImport.MintKey appends _2, _3 and on when the fallback is itself taken — the key is unique before it is pretty. The definition takes the canvas name, truncated to 255 characters (CanvasWorkflowImport.MaxNameLength, which is the name column's own bound — GraphWorkflowDefinitionConfiguration declares HasMaxLength(255)), and the description "Imported from Open Canvas (canvas workflow {id})." — provenance without a schema change. The create and update validators cap a name at 200 (GraphWorkflowRequestLimits.MaxNameLength), so an imported name of 201 to 255 characters lands in the database but has to be shortened before the definition can be saved again.

The importer emits no positions. position is optional and the editor already lays out a node that arrives without one, so generating a second layout rule here would leave one of the two dead. The practical consequence: an imported definition opens laid out and unsaved. The canvas is dirty the moment you open it, and the layout is persisted when you save. That is expected, not a bug.

The context edge around a Pause. Open Canvas's Pause was a pass-through resume: its post-adapter forwarded the answer it was waiting on unchanged. A Graph Workflow Pause writes a decision document of its own — {decision, comment, payload} — and a node's input is its ONE satisfied predecessor's output document, becoming the upstream map only when there are several (§3.2). Mapped one-for-one, Agent A → Pause → Agent B therefore hands B the approval metadata and never A's answer, and a Pause before an End loses the result the same way. That was observed live in the S4 round, not reasoned about.

So the importer emits one extra unconditional edge from the pause's nearest non-Pause ancestor X to each starved successor Y. Y keeps the default All join policy, so it is admitted only once both the content edge and the pause's own approved edge are satisfied — never ahead of the approval — and with two satisfied predecessors its input is the upstream map { "<X>": …, "<P>": … }. An imported Agent carries includeUpstreamOutputs: true and so sees X's text; an imported End carries resultPath: null and so keeps both documents. The rule applies to every pause and walks back through consecutive ones, so A → P1 → P2 → B gains both A → P2 and A → B; a successor several pauses reach is judged once, over the union of what all of them are fed by. X may be the Start node, whose output is the run's own input, which is exactly the content the pause interrupted.

The decision is keyed on the successor, and the guards are the validator's own (GraphWorkflowGraph.PauseContextWarnings), read off the mapped node's joinPolicy and kind rather than off a canvas kind — an importer that advised an edge the validator refuses would produce a definition nobody can save again. A successor is owed an edge only when it is starved (every one of its inbound edges leaves a Pause; one something else already feeds loses nothing), its joinPolicy is not Any (of any kind — an unconditional content edge would admit an Any node ahead of every approval, including when all of them were rejected), and its pauses' nearest non-Pause ancestor is unique and is not a Condition (two candidates are mutually exclusive branches, and an edge from each would hang an All successor on the branch never taken; a Condition would gain a second unconditional out-edge, which §2.4 refuses). Three more cases add nothing: a pause with no successor (already an IMPORT NEEDS ATTENTION: graph, §9.3), a pair the canvas already wires — a second unconditional edge over one pair is a validation error (§2.4) — and a self-loop.

Every guard but the starvation test is unreachable for a real import: an Open Canvas graph carries no Condition, no Join and no join policy, and every stored shape is linear. They are stated in code anyway because the rule, not today's canvas vocabulary, is what the next node kind has to keep holding.

The editor offers the same edge to an author, on the connect gesture (§4.6), under the same three guards. The one difference is when it runs: the editor runs only on that gesture, so an author who deletes the edge keeps it deleted, while the importer runs once over a whole canvas.

maxAttempts: 1 on an imported Agent is deliberately below the default of 3. An import is conservative: re-running somebody's agent turn twice more, on a graph they have not looked at since it changed shape, is not a decision this conversion gets to make for them.

9.3 IMPORT NEEDS ATTENTION:

The mapper is total — it always returns a document and never refuses a graph. Whatever it could not translate faithfully becomes a reason.

A converted graph then goes through IGraphWorkflowDefinitionService, which owns the parse, the node cap and the hash and count. If that validation refuses it, the definition is saved anyway, through IGraphWorkflowStore and past the validator, with its description prefixed:

IMPORT NEEDS ATTENTION: <reason>

Such a definition cannot be run until an operator opens it in the editor and fixes what the validation strip names. Nothing is discarded for being invalid. A Debug node with no successor, a shape the new validator refuses, an unknown kind — all of them arrive as a definition you can see and edit, rather than as an absence you have to notice.

A blob that does not decrypt or parse as JSON is reported at Error with its canvas id and exception type, then the read fails and startup stops before destructive migrations. The source and recovery ciphertext remain available for repair or restoration with the original node key. A whole-read or staging failure also stops startup; it is never treated as an empty successful read.

The write half records mapping or persistence failures at Error and refuses to commit any definitions when one row failed. Validation refusals still use the needs-attention path and preserve the graph. Only after every candidate has been preserved does the transaction remove recovery staging and commit. Startup propagates failures instead of serving a node whose conversion silently lost data.

Graph content is never logged. Instructions and start text are exactly the payload the column is encrypted to protect. The reader logs entry and per-row failures; the writer logs one line per clean import. The summary an operator must not miss is logged at Warning, only after the import and recovery cleanup commit:

Open Canvas one-shot import complete: {Imported} imported, {NeedsAttention} need attention, {Failed} failed.
Open Canvas has been removed; canvas_workflows is dropped.

One further Warning line follows per needs-attention canvas, naming the id, the name and the reason. A failed canvas gets an Error line instead, and it names no reason: the read-half line carries the id and the exception type, the write-half lines the id, the name and the exception type. A cleanly imported canvas whose mapping still changed something gets one Warning per change.

9.4 Recovery and backups

The same-database encrypted staging is the recovery mechanism for this conversion. The existing INodeDbBackupService.BackupBeforeMigrationAsync() backup is supplementary: its best-effort failure no longer removes the only recovery copy. If the durable staging write fails, startup stops before migrations. If import fails after the source table has been dropped, staging remains and the next startup retries it.

Back up the complete data directory, including its key material, before upgrading. This fix cannot recover canvases that an earlier version already dropped without importing; those require a pre-upgrade backup. Do not delete or edit the recovery table to bypass a startup failure. Resolve the reported underlying fault or restore a complete backup. No plaintext export is created.


Related pages