Skip to content

fix(coordination): make parallel output composition deterministic #720

Description

@proffesor-for-testing

Problem

WorkflowOrchestrator treats ready steps without dependsOn as a parallel group and executes them with Promise.allSettled (exact main, lines 470-489). Each asynchronous step maps its output into the shared workflow context before the group joins (lines 568-573), and mapStepOutput writes arbitrary outputMapping target paths directly into that shared object (lines 681-687).

The YAML loader validates only that each step's mapping is a string-to-string record (lines 199-205); it does not compare write sets across concurrently runnable steps. Two independent steps may therefore map to the same path, or to ancestor/descendant paths such as analysis and analysis.security, with the final state determined by completion order. Individual step success does not prove that the composed workflow state is valid.

This is exact-head static evidence, not a runtime reproduction. It is nevertheless a concrete unsafe composition path: the shared-state mutation occurs before the parallel barrier, there is no conflict detector or declared merge policy, and no combined-state postcondition is evaluated.

The motivating research is Passes Alone, Fails Together (arXiv:2609.25396). It found almost no naturally mined interference after correcting its grader—1 case in 834 runs over 417 Django change pairs—but 97% interference in deliberately constructed shared-helper tasks. A short coordination message recovered 82% of those constructed failures. The constructed rate must not be read as real-world prevalence; the useful result is the controlled demonstration that individually valid changes are not compositional evidence.

Proposed solution

P0 — isolated outputs and fail-closed composition

  1. During a parallel group, capture each step's raw result and mapped output in a step-local immutable record. Do not mutate shared context.results or arbitrary context paths until all steps in the group reach a terminal disposition.
  2. Derive a normalized write set from every outputMapping target.
  3. Detect exact-path and ancestor/descendant conflicts before shared-state mutation.
  4. Reject conflicting writes by default with a typed parallel_output_conflict result. Do not select a winner by array order or completion timing.
  5. Compose disjoint writes in a stable order, then validate the combined state before committing it atomically.
  6. Emit a versioned ParallelCompositionReceipt containing workflow/revision identity, parallel-group and step IDs, normalized write sets, step-output hashes, detected conflicts, declared strategy, combined-state hash, and disposition (composed, conflict, partial, or invalid).

P1 — explicit joins and deterministic reducers

  1. Support an explicit join/integration step that consumes named step outputs instead of implicit shared writes.
  2. If overlapping writes are intentionally allowed, require a declared reducer/merge strategy with input/output schemas. The reducer must be deterministic, side-effect-free, and independently testable.
  3. Evaluate group-level postconditions after composition. A group is not successful merely because every member step returned success.
  4. Treat reads of another concurrently runnable step's output as a dependency error; require dependsOn rather than relying on timing.
  5. Preserve failed, skipped, cancelled, and unavailable step outputs as typed evidence without partially applying their mappings.

P2 — semantic-coordination qualification

Add a small controlled corpus of workflows where steps pass independently but can conflict through shared paths, shared schemas, or integration invariants. Compare no-message, explicit dependency, explicit join, and declared-reducer conditions. Keep constructed-task results separate from estimates of field prevalence.

Suggested types

interface ParallelCompositionReceipt {
  version: 1;
  workflowId: string;
  workflowRevision: string;
  executionId: string;
  groupId: string;
  steps: Array<{
    stepId: string;
    disposition: 'succeeded' | 'failed' | 'skipped' | 'cancelled';
    outputHash?: string;
    writes: string[];
  }>;
  conflicts: Array<{
    pathA: string;
    stepA: string;
    pathB: string;
    stepB: string;
    kind: 'exact' | 'ancestor-descendant';
  }>;
  strategy: 'disjoint' | 'explicit-join' | 'declared-reducer' | 'rejected';
  combinedStateHash?: string;
  disposition: 'composed' | 'conflict' | 'partial' | 'invalid';
}

Acceptance criteria

  • Parallel step results remain isolated until the group reaches its barrier.
  • Exact target-path collisions are detected before shared-state mutation.
  • Ancestor/descendant collisions are also detected.
  • Conflicts fail closed unless a declared reducer or explicit join owns them.
  • Disjoint mappings preserve current workflow behavior.
  • Declared reducers are deterministic, schema-validated, and receipt-bound.
  • Reversing step completion order or injecting delays yields the same final-state hash and disposition.
  • A failed/cancelled step cannot leave partially applied shared mappings.
  • A combined-state postcondition can reject individually successful steps.
  • Cross-step reads without dependsOn are rejected or surfaced as invalid.
  • YAML-loaded and programmatic workflows apply the same composition rules.
  • CLI/MCP/persistence output exposes the composition receipt and conflict details without leaking sensitive raw output.

Validation experiments

  1. Delay inversion: run two conflicting steps with reversed artificial delays; both executions must reject with equivalent receipts.
  2. Exact collision: map two outputs to context.summary; verify zero silent overwrites.
  3. Prefix collision: map to context.analysis and context.analysis.security; verify conflict classification.
  4. Disjoint control: map to unrelated paths; verify deterministic composition and no regression.
  5. Reducer order: permute inputs repeatedly; either prove the declared order is part of the contract or require commutative/idempotent behavior.
  6. Partial failure: one step succeeds and one fails/cancels; verify no partial group commit.
  7. Stale read: let one parallel step attempt to consume another's output; verify that a dependency is required.
  8. Semantic integration: individually schema-valid outputs violate a group invariant; verify the postcondition rejects the group.
  9. Race stress: repeat with randomized delays and worker counts; require one stable receipt/final-state hash.
  10. Compatibility: run existing built-in workflows with disjoint mappings and compare behavior and persisted output.

Relationships and non-goals

Research

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions