Skip to content

Latest commit

 

History

History
221 lines (186 loc) · 12 KB

File metadata and controls

221 lines (186 loc) · 12 KB

koru fleet — one supervisor for every koru-managed project

koru autonomous up drives exactly one project (--project). Running it for N projects meant hand-configuring N systemd units, each independently subject to systemctl stop/restarts, with no single place to see "is koru working on anything right now?" across the whole machine.

koru fleet up (src/koru/cli_fleet.py) is a thin supervisor: it discovers every project that opted into koru's LLM-agent policy (.planfile/.koru/policy.yaml, written by koru --init) under a workspace root, and runs one supervised koru autonomous up child process per project — a single broker-style service coordinating many project "topics" (one child per project), analogous to how an MQTT broker manages many topic subscribers from one process.

Why process-per-project, not thread-per-project

Each project's autonomous loop keeps its own crash/resource blast radius — matching how --replace-existing / --allow-duplicate already reason about one loop per project. A runaway or crashing project can't take down every other project's loop. The tradeoff is supervisor-level bookkeeping (start, poll, restart-with-backoff, terminate-on-shutdown) instead of just spawning threads; cli_fleet.py keeps that bookkeeping in one small module (_ManagedProject).

Architecture

flowchart TB
    subgraph Fleet["koru fleet up (systemd: koru-autonomous.service)"]
        D[discover_projects workspace] --> M[_ManagedProject per project]
        M --> P1["autonomous up --project semcod --replace-existing"]
        M --> P2["autonomous up --project koru --replace-existing"]
        M --> P3["autonomous up --project goal --replace-existing"]
        M --> P4["... one child per discovered project"]
        L["poll loop (every 2s)\nrestart on exit + backoff\nrescan every 300s"] -.supervises.-> P1
        L -.supervises.-> P2
        L -.supervises.-> P3
        L -.supervises.-> P4
    end
    D -->|".planfile/.koru/policy.yaml\nmarks a koru-managed project"| FS[("~/github/** filesystem")]
Loading

ASCII view of the same shape, for a terminal/no-mermaid-renderer read:

                     ┌─────────────────────────────┐
                     │   koru fleet up (1 process)  │
                     │   systemd: koru-autonomous    │
                     └───────────────┬───────────────┘
                                     │ discover_projects(~/github)
                                     │ (finds .planfile/.koru/policy.yaml)
                 ┌───────────────────┼───────────────────┬─────────────────┐
                 ▼                   ▼                   ▼                 ▼
        ┌────────────────┐  ┌────────────────┐  ┌────────────────┐  ┌───────────┐
        │ autonomous up   │  │ autonomous up   │  │ autonomous up   │  │   ...     │
        │ --project semcod│  │ --project koru  │  │ --project goal  │  │  (N more) │
        │ --replace-exist.│  │ --replace-exist.│  │ --replace-exist.│  │           │
        └────────────────┘  └────────────────┘  └────────────────┘  └───────────┘
              child pid            child pid            child pid       child pid

        poll every 2s -> still alive? keep going. exited? restart after backoff.
        rescan every 300s -> new `koru --init`-ed project appears? add it, no restart needed.
        SIGTERM/SIGINT to the fleet -> terminate() every child, wait up to 30s, then kill -9.

Usage

koru fleet ls                                    # preview discovery, no processes started
koru fleet ls --workspace /path/to/workspace-root

# Bootstrap many sibling repos under one parent folder (idempotent, no --force)
koru fleet bootstrap ~/github/subactor --dry-run
koru fleet bootstrap ~/github/subactor --umbrella \
  --include runtime --include core --include agents \
  --exclude backups --exclude logo

koru fleet up                                    # supervise every discovered project, ~/github default
koru fleet up --workspace /path/to/root \
  --restart-backoff-seconds 30 \
  --rescan-interval-seconds 300 \
  -- --ide claude --ticket-sources all            # everything after `--` is forwarded to each child

koru fleet bootstrap (alias: koru fleet init)

Takes a parent directory (workspace folder), discovers child projects (directories with .git, optionally filtered by --include / --exclude), and ensures each has .planfile/ + .planfile/.koru/policy.yaml so koru fleet ls can see them.

Flag Default Effect
workspace (positional) $KORU_FLEET_WORKSPACE or ~/github Parent folder to scan
--umbrella off Also initialise the workspace root itself (git optional)
--dry-run off Report actions without writing
--include GLOB (all) Only match basename/relative path (repeatable)
--exclude GLOB + built-in backups, node_modules, … Skip matches (repeatable)
--depth N 1 Walk depth for nested repos
--allow-non-git off Consider dirs without .git
--force off DANGEROUS — re-runs koru --init --force (overwrites tickets; writes .bak-*). Never needed just to add a missing policy.yaml

Soft ensure (no clobber): if a child already has .planfile/config.yaml but is missing policy.yaml, bootstrap writes only the policy stub (+ .gitignore entry). It does not import starter tickets and does not require --force. This is the fix for the old "had to --init --force to get fleet coverage and lost all tickets" failure mode.

MCP: there is no fleet MCP tool yet — use the CLI (koru fleet bootstrap / koru fleet ls). Planfile MCP tools (koru_list_tickets, …) remain per-project and take project_root.

koru fleet up / koru fleet ls flags

Flag Default Effect
--workspace $KORU_FLEET_WORKSPACE or ~/github Root to discover koru-managed projects under
--restart-backoff-seconds 30 Delay before restarting a project's loop after it exits
--rescan-interval-seconds 300 How often to re-discover projects, so a newly koru --init-ed / bootstrapped project joins without a fleet restart
-- <args> Forwarded verbatim to every koru autonomous up child (e.g. --ide, --ticket-sources)

Discovery prunes obvious non-project noise during the walk (test-data, tests, examples, plugins, archive, rebuild, node_modules, VCS/venv dirs) — see _JUNK_PATH_SEGMENTS in src/koru/cli_fleet.py. A project nested inside another koru-managed project (e.g. semcod/koru inside semcod) is legitimate and gets its own loop; only known junk directory names are excluded, not nesting itself.

Deploying as a systemd user service

Copy examples/systemd/koru-fleet.service.example to ~/.config/systemd/user/koru-autonomous.service, adjust the paths, then:

systemctl --user daemon-reload
systemctl --user enable --now koru-autonomous.service
systemctl --user status koru-autonomous.service --no-pager

Restart=always on the fleet unit only needs to cover the supervisor process crashing outright — each project child already has its own restart-with-backoff handled inside koru fleet up.

The bug this replaced a hand-rolled fix for

Building and load-testing this fleet surfaced a real, previously-untested bug in the existing --replace-existing process-matching logic (autonomy/operator/operator_processes.py), reproduced live in this session:

sequenceDiagram
    participant OldProc as Old process (real)<br/>cwd=~/github/semcod<br/>cmd: --project .
    participant FleetChild as New fleet child (test)<br/>cwd=/tmp/fleet-test/proj-a<br/>cmd: --project .
    Note over OldProc: Running for hours, healthy
    FleetChild->>FleetChild: _command_project("--project .")
    FleetChild->>FleetChild: resolves "." against **its own** cwd<br/>(BUG: should resolve against<br/>OldProc's cwd instead)
    FleetChild->>OldProc: "your --project . equals MY project path!"
    Note over FleetChild,OldProc: False match: two unrelated<br/>relative "--project ." processes<br/>collapse onto the same path
    FleetChild-->>OldProc: --replace-existing kills it
    Note over OldProc: Dead. Unrelated project's<br/>hours-long loop lost for no reason.
Loading

Any two koru autonomous up --project . --replace-existing instances anywhere on the machine — not just deliberately-concurrent fleet children — were at risk of this, since --project . (relative) is the invocation shown by koru --doctor's own recovery hint. This is a strong candidate for at least some of the autonomous loop's previously observed "why doesn't it stay running" unreliability.

Fix: resolve a relative --project value against that process's own cwd (_process_cwd(pid), already computed by the caller) instead of the checking process's Path.cwd(). See tests/test_operator_processes_project_matching.py for the regression coverage (11 tests, including a direct reproduction of the two-unrelated-instances collision) — this function had zero prior test coverage.

A second, related race was also fixed in the same investigation: a concurrent actor (a human, or another koru/agent session) closing a ticket while a tillm_shell-driven vendor CLI (claude -p ..., can take minutes) was still mid-flight could get its finished work reopened by a stale shell_drive_finalize verify run. See post-run-verify.md for the general queue.post_run_verify mechanism this interacts with, and tests/test_shell_drive_finalize.py for the _ticket_already_resolved() fix.

Known limitations / next steps

  • No shared dashboard yet across fleet children — koru serve --workspace already supports multi-project discovery for the read-only dashboard; wiring koru fleet up to also launch (or point at) one shared koru serve --workspace instance instead of N per-project ones is a natural follow-up.
  • --rescan-interval-seconds only adds newly discovered projects; a project that stops matching the policy marker (e.g. .planfile/ removed) is not currently removed from the managed set until the fleet restarts.
  • No per-project resource caps (CPU/memory) — a single project's heavy scan can still slow down the machine for all sibling children, even though it can no longer kill them.
  • _ManagedProject.command() resolves the koru binary via sys.argv[0]; this is correct for the common case (systemd ExecStart uses an absolute path) but would fall back to a bare "koru" on $PATH if koru fleet up were ever invoked through a wrapper that rewrites argv[0].

See also