Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .cursor/skills/orchestrator-executor/SKILL.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,6 @@
---
name: orchestrator-executor
description: Fable 5 orchestrator-executor pattern for big tasks
---

You (Fable) are the orchestrator. Plan, decompose, synthesize. Reasoning-heavy phases go to deep-reasoner (5.6 Terra). Mechanical work goes to fast-worker (5.6 Luna). For high-stakes decisions, run deep-reasoner twice with slightly different framings and synthesize the best of both. Keep your own context lean. Delegate rather than doing mechanical work yourself.
78 changes: 78 additions & 0 deletions docs/effort-graph/CONTEXT.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,78 @@
# Effort Graph glossary — reasoning primitives for `@flatbread/proof`

This glossary defines the **epistemic primitives** that make up the Effort Graph — a persistent, queryable memory layer over the reasoning and planning that happens during long-horizon software work, single-agent or multi-agent.

The Effort Graph is **built on top of** Flatbread's content-layer vocabulary (see [Flatbread glossary](../glossary.md) for `Collection`, `Record`, `Relation`, `Refs`, `ID`). Each primitive here is a Flatbread **Collection**; instances are **Records**; cross-primitive references are **Relations** wired through frontmatter `refs`.

**What this is not:** a CMS, an authoring UI, a hosted memory product, or a general task tracker. It is a relational substrate for capturing the **gray-area reasoning** that an ADR-only record loses: open questions, considered alternatives, sticky constraints, prospective risks, and post-hoc invalidations.

**Operational provenance** (which session produced this, which agent, which model, which DAG run) is captured as **frontmatter fields** on these primitives, not as peer collections. The durable transcript record lives next to the graph under `.flatbread/artifacts/` (see [`packages/proof/README.md`](../../packages/proof/README.md) §Artifact Output).

---

### Effort

The **anchor** of the graph. One Effort represents a coherent, named thread of work — a feature, a migration, a spike, a research investigation, a refactor. Every epistemic primitive belongs to exactly one Effort.

An Effort is the stable filter (`effort: { eq: "<effort-id>" }`) that scopes every "what's still open / what did we conclude / what are we considering?" query. Loss of the Effort anchor is the failure mode that vault MCPs and flat memory stores cannot avoid; preserving it is the central wedge.

An Effort has its own lifecycle (active, paused, completed, abandoned) but carries **no reasoning content of its own** — its body is a short description; the reasoning lives in the primitives that ref back to it.

### Issue

A **tracked unit needing attention within an Effort**, in the GitHub-issue sense — broader than "something is wrong." Issues span open questions, observed defects, identified gaps, and explicit blockers. Each Issue carries a `kind` field that names the speech act (`question`, `defect`, `gap`, `blocker`, …) and a status (`open`, `resolved`, `deferred`, `wontfix`).

An Issue is resolved by a Decision (we'll do X) and/or one or more Findings (here's what we learned that closes this). The `kind` is open-ended (free-form string) so common values emerge from dogfooding rather than from schema enforcement.

Feature _proposals_ are not Issues — they are `Decision{state: proposed}`. Issues are reactive (something exists that needs attention); proposed Decisions are proactive (let's commit to doing X).

### Finding

A **grounded observation** — a claim about reality (the codebase, the user, the literature, the runtime) backed by cited evidence. Findings resolve Issues, support or contradict Decisions, surface Risks, and invalidate prior Findings or Decisions when reality refutes a prior belief.

The `Finding{kind: retrospective}` variant carries the additional semantic that the Finding was produced **after a Decision shipped** and may invalidate that Decision in light of new evidence. Other Finding kinds (e.g. `measurement`, `survey`, `dead-end`) may emerge from usage but are not load-bearing in the schema.

### Decision

A **commitment** — a chosen path among alternatives. Has a `state`: `proposed` (under consideration), `accepted` (committed), `rejected` (an alternative we chose not to take), `superseded` (replaced by a later Decision), or `deprecated` (no longer current but not replaced).

Multiple `state: proposed` Decisions under the same Effort represent **competing directions under exploration**. When one is `accepted`, the others should transition to `rejected` with a back-pointer to the accepted Decision. This is the schema's substitute for a separate `Proposal` primitive.

A Decision cites the Findings, Constraints, and Risks it weighed; it does not duplicate their content.

### Constraint

A **sticky boundary** that scopes the decision space for an Effort. May be hard (license incompatibility, regulatory rule, irreversible upstream choice) or soft (team preference, budget envelope, performance target). Constraints typically outlive individual Decisions and apply to many of them.

A Constraint is not a Risk: a Constraint is a known limit you must design within; a Risk is a possible outcome you might suffer.

### Risk

A **prospective negative outcome** with a likelihood and a severity. Risks attach to Decisions as part of the rationale for choosing among them. A Risk has a lifecycle: `open` (live, unmitigated), `mitigated` (an accepted Decision exists to reduce likelihood or severity), `realized` (it happened — usually triggers a Finding and possibly a retrospective Finding), or `accepted` (we knowingly proceed despite it).

---

## Cross-cutting edge vocabulary

These edges are **ubiquitous** — they live on every epistemic primitive. They are the type-agnostic semantic graph that lets a reader trace causality, evolution, and disagreement.

| Edge | Description |
| -------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `derives_from` | Causal upstream — what this artifact is responding to or built on. (A Finding `derives_from` an Issue; a Decision `derives_from` Findings + Constraints; a retrospective Finding `derives_from` the original Decision.) |
| `supersedes` / `superseded_by` | Replaces an earlier artifact of the same primitive. The forward edge (`supersedes`) is canonical; `superseded_by` is a derived projection materialized to disk so any single record can answer "am I current?" in one access (see ADR-0004). |
| `invalidates` / `invalidated_by` | Stronger than `supersedes` — asserts the targeted artifact was _wrong_, not just outdated. Used primarily by retrospective Findings against shipped Decisions. Same forward-canonical / materialized-back-edge rule as `supersedes`. |

**The edge vocabulary is permitted to grow** as dogfooding surfaces real omissions. Candidate additions to watch for: `refines` (soft non-replacing evolution), `contradicts` (explicit disagreement that doesn't yet rise to invalidation), `blocks` (an open Issue gating progress on another). Any addition must justify itself with a query the existing vocabulary cannot answer.

---

## What is intentionally not modeled

- **Session, Run, Plan, Artifact, Agent** as collections. These are operational provenance, captured as opaque-string frontmatter fields (`produced_in`, `created_by`, etc.) on the epistemic primitives above. Their durable log-grade record lives under `.flatbread/artifacts/` from `@flatbread/proof` runs.
- **Investigation** as a collection. An investigation is a Session-grouping of Findings (and possibly an Issue with `status: investigating`), not a noun in its own right.
- **Question** as a collection. Collapsed into `Issue{kind: question}` — the speech-act distinction does not warrant a separate primitive.
- **Proposal** as a collection. A Proposal is a `Decision{state: proposed}`.
- **Retrospective** as a collection. A Retrospective is a `Finding{kind: retrospective}`.
- **Branch** as a frontmatter field. Speculative exploration lives on git branches; cross-branch reasoning is preserved by promoting artifacts to the integration branch when an exploration closes (rejected or merged).

These collapses may be revisited if real usage proves the host primitive cannot carry the missing semantics; Flatbread's `refs` model permits later splitting without ID breakage.
28 changes: 28 additions & 0 deletions docs/effort-graph/adr/0001-effort-graph-memory-location.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
# 0001 — Effort Graph memory location

Status: Accepted

## Context

The Effort Graph stores epistemic artifacts (Issues, Findings, Decisions, Constraints, Risks) as flat markdown files indexed by Flatbread. Where those files live relative to the project repo determines whether reasoning branches with the code, and whether reasoning survives when a speculative branch is abandoned.

Three modes were considered:

- **A. In-repo, branch-coupled** — files under `<project>/.flatbread-efforts/`. Reasoning branches with the code. Abandoned branch ⇒ abandoned reasoning.
- **B. In-repo with promote-on-close tooling** — same location, plus a helper that lifts artifacts onto the integration branch when an exploration closes. Requires a robust definition of "close a branch" across squash-merge, rebase-merge, PR-closed-unmerged, branch-deleted-then-resurrected.
- **C. Sibling repo / submodule** — memory has its own git history independent of project branches. Cross-branch reasoning loss disappears because memory commits to memory `main`.

A cross-branch `ref` mechanism (point a ref at reasoning on another branch) was rejected: refs must resolve at index time against a single on-disk tree; resolving across branches requires either mutating the working tree or a per-branch index (a source-plugin rewrite), breaks the self-contained-artifact review story, and makes staleness invisible on rebase/force-push/delete.

## Decision

Ship **mode A as the default**. Make the schema and write API **identical across all three modes** so that **mode C works by repointing `path` in `flatbread.config.ts`** with no code change; document mode C as the supported alternative for teams that abandon exploration often or span multiple repos.

**Defer mode B.** The promote-on-close ergonomics can be replaced day one by a documented `git cherry-pick <subdir>` + status-flip convention (`Decision: rejected_explored`, `Issue: wontfix`, `Finding: archived-from-exploration`). A `flatbread efforts promote-branch-artifacts` helper may land later.

## Consequences

- The default (A) accepts that reasoning on a never-merged branch is lost unless the team promotes it. This is acceptable because most efforts merge.
- Teams that care about preserving rejected exploration have two escape hatches without new platform code: cherry-pick promotion (still mode A) or mode C (config-only).
- The schema must not encode git internals (no `branch:`, no `git_ref:` field). Branching is emergent from git itself plus `Decision.state` and `derives_from` edges.
- Deferring B leaves a documented manual convention as the only path for promote-on-close until demand justifies the helper.
28 changes: 28 additions & 0 deletions docs/effort-graph/adr/0002-semantic-mutation-write-surface.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
# 0002 — Semantic-mutation write surface

Status: Accepted

## Context

Storing edges bidirectionally (`supersedes`/`superseded_by`, `invalidates`/`invalidated_by`) so that any single record can answer "am I current?" in one access (ADR-adjacent decision captured in `CONTEXT.md`) means a single conceptual change — "Decision X supersedes Decision A" — must update **two files atomically**: write X with `supersedes: [A]`, and patch A with `superseded_by: X`. A half-applied edge (X claims supersession, A does not know) is a corruption mode.

Two stances for the write API the agent calls:

- **(α) Narrow create-only surface.** Ship only `WriteIssue` / `WriteFinding` / `WriteDecision` / `WriteConstraint` / `WriteRisk` create mutations plus a generic frontmatter patch. The agent is responsible for issuing both halves of a bidirectional edge; the read shim warns about half-applied edges during a validation pass.
- **(β) Semantic-mutation surface.** Each schema-level concept (supersede, invalidate, resolve-issue, mitigate-risk) gets a dedicated typed mutation that atomically updates both ends of the edge. The platform owns bidirectional consistency; the agent does not.

## Decision

Adopt **β — a semantic-mutation write surface**. Mutations are validated and expanded at the level of **semantic edits**, not file edits. Each mutation:

- has its own Zod schema and validation rules (e.g. `SupersedeDecision` requires the target Decision to exist in the index and not already be `superseded`),
- expands into the full set of file writes its semantics imply, and
- completes all of those writes or none (transaction semantics — partial-failure handling is specified separately).

## Consequences

- The agent-facing write contract is a set of named mutations, not "validate this YAML and write it." This is a larger design and implementation surface than α.
- Bidirectional-edge integrity is guaranteed by the platform, mirroring how `@flatbread/proof` put convergence-loop semantics in the DAG primitive rather than in agent prompts.
- A transaction/rollback primitive is now required (what happens when the second of two writes fails). This is a new capability for the write path.
- Flatbread has no mutation support today; β raises the question of whether mutations are exposed through GraphQL resolvers in core or through a standalone write library that does not touch the read layer. That fork is decided separately.
- The set of semantic mutations must be enumerated and kept small; each new mutation is API surface that must be versioned and taught to agents.
Original file line number Diff line number Diff line change
@@ -0,0 +1,37 @@
# 0003 — Write-path architecture and read-after-write consistency

Status: Accepted

## Context

ADR-0002 adopted a semantic-mutation write surface (β). Flatbread today is read-only and in-memory: `FlatbreadProvider` (`packages/core/src/providers/base.ts`) builds the GraphQL schema once at construction and exposes only `query()`. There is no `Mutation` type in `packages/core` and no write-back. The filesystem is the source of truth; the GraphQL graph is a projection built once. Live reload is explicitly unsupported today ([issue #65](https://github.com/FlatbreadLabs/flatbread/issues/65); `docs/positioning.md`).

Two write-path architectures were considered:

- **(1) Mutations inside core's GraphQL.** Add a `Mutation` type and resolvers that write files; the provider gains `mutate()`. Forces an in-process cache-coherence subsystem (every successful write must patch/invalidate the cached `EntryNode` graph or the next `query()` is stale) into core, which it does not have today.
- **(2) Standalone writer + read-only GraphQL.** A separate library owns the Zod mutation schemas, file expansion, and transaction semantics, and writes the source-of-truth files directly. The GraphQL read layer stays read-only and re-projects from disk.

A separate question is the **read-after-write consistency model**: tool-call-boundary re-index (write returns touched ids/paths; next read re-indexes) vs instantaneous in-process read-after-write.

## Decision

Adopt **architecture (2): a standalone semantic writer; GraphQL stays read-only.** Mutation logic (validation, multi-file expansion, transactions) lives outside core's resolver layer, consistent with Flatbread's existing files-are-source-of-truth model. A thin GraphQL mutation _facade_ that delegates to the writer may land later, but GraphQL mutations are not the home of the logic.

Adopt **live-reindex (watch mode) as a v1 dependency** for the consistency model, satisfied by implementing the **"Draft unified watch design"** already specified in `docs/local-dev-loop.md` (reload records → re-run ID/ref/cardinality validation → rebuild schema → hot-swap the live schema only if the new graph validates, else keep the prior schema and log).

### Consistency contract (v1)

The in-memory graph that reads are served from is rebuilt after a write; there is a brief window between "file saved" and "rebuild complete." During that window a read does not block and never returns a torn/half-built graph, but a _concurrent_ reader may observe the prior graph (briefly out of date, never wrong-shaped). The window widens with corpus size because rebuilds are full today. The contract that bounds this:

1. **Read-your-own-writes (mandatory).** A mutation's return payload includes the written/changed artifacts, so the writing agent never re-queries to see its own write. This removes the window entirely for single-agent write-then-read, which is the dominant case.
2. **Default eventual, opt-in strict for concurrent readers (Q7.i → Option 1).** A second, concurrent reader may be a beat behind by default. When a workflow cannot tolerate this (e.g. a downstream `@flatbread/proof` task that must observe an upstream task's writes), the reader opts into a strong read that waits until the index has caught up to the depended-on write before answering. **This relaxed-by-default behavior, and how to request a strict read, must be called out in user-facing documentation when implemented.**
3. **Incremental reindex (v1 dependency, Q7.ii → Option A).** Because the writer knows exactly which files it touched, reindex re-reads only the changed files plus their ref-affected neighbors and patches the in-memory graph, instead of rebuilding the whole graph. This keeps the staleness window small regardless of corpus size (thousands+ of memories) and pairs with retiring the per-resolver `cloneDeep(contentNodesByCollection[...])` cost in `packages/core/src/generators/schema.ts`.

## Consequences

- The Effort Graph spec now has a **hard dependency on shipping the unified watch / live schema-swap seam** in Flatbread, **including incremental (changed-files-only) reindex**. This is scoped (a documented design contract and an existing codegen watch loop to factor from), not greenfield, but it is on the critical path and must be sequenced before the write story is considered done.
- Writes cannot destabilize core's read path, since transactional file-writing lives in a separate package.
- The writer must return the ids and file paths it touched, both for read-your-own-writes payloads and so the incremental reindex layer can refresh exactly the affected collections.
- A reader opting into a strict read needs a way to name the write generation it depends on; the writer must therefore expose a monotonic generation/version token in its return payload.
- Transaction/rollback semantics for multi-file mutations remain to be specified (see follow-up).
- If the watch seam slips, the fallback is tool-call-boundary re-index (read shim re-indexes affected collections per invocation); this is a degraded mode, not the target.
Loading
Loading