Skip to content

No error explanation if a story run fails #377

Description

@andreykeycee

What you were doing

I set up facility for my project and created a story for an architect agent to plan the refactoring

Where it went wrong, or slower than expected

I've got an error after 3 minutes of running, no evidence of what exactly went wrong, only "claude_code exited with status 1".
After a while I researched that the agent was failing with authentication due to wrong API token

What you expected instead

Clear stacktrace/error message in UI so I understand what exactly went wrong so I can adjust from there

Time from clone to first agent run

2.5 hours

Activity

  1. changed the title [-]Engine events are discarded when a turn fails, hiding the real error[/-] [+]No error explanation if a story run fails[/+] on Sep 14, 2026
  2. andreykeycee commented on Sep 14, 2026

    @andreykeycee
    ContributorAuthor

    Adding the root cause, since I had the source open while debugging this.

    Version: main @ 771b647 (0.12.0) · local source install, FACILITY_WORKSPACE_DRIVER=docker

    What the error actually was

    401 API key is invalid — a bad value in the project's managed ANTHROPIC_API_KEY. Recovering
    that meant finding the workspace container and grepping the Claude Code session transcript at
    /workspace/.facility/claude/projects/<slug>/<session>.jsonl. Nothing in the UI, the story
    timeline, or turn_events held it: turn_events had exactly two rows, turn.started and
    turn.failed, the latter carrying only claude_code exited with status 1.

    Root cause

    The mechanism to surface this already exists and is correct — it is just scoped to one error
    code.

    CliAgentEngine.execute attaches the parsed engine events to every failure it throws
    (services/api/src/turns/engines.ts:93-101), for both agent_session_corrupt and
    agent_engine_failed. TurnDispatcher.persistFailedEngineEvents
    (services/api/src/turns/dispatcher.ts:536-551) writes those events with redaction and bounding,
    exactly like the success path.

    But it has a single call site, dispatcher.ts:222, inside this guard
    (dispatcher.ts:215-221):

    } catch (error) {
      if (
        !(error instanceof AgentEngineError) ||
        error.code !== "agent_session_corrupt" ||
        !session
      ) {
        throw error;
      }
      await this.persistFailedEngineEvents(error, eventBase, secrets);

    So a resumed-session failure gets its events persisted, and every other engine failure rethrows
    at line 220 without them
    . An agent_engine_failed — which is what an API error such as a 401
    produces — reaches the outer catch at dispatcher.ts:333-354, which records cost from
    error.details.usage and writes one turn.failed carrying only the message string.
    error.details.events is dropped on the floor.

    turn-dispatcher.integration.test.ts:506 already proves the mechanism works and redacts correctly
    on the corrupt-session path, so this is a narrowing bug rather than a missing feature.

    Why the message is always the generic fallback. engines.ts:87:

    const message = result.stderr.trim() || `${this.name} exited with status ${result.exitCode}`;

    The Claude Code CLI reports API errors as a result object on stdout with is_error: true,
    leaving stderr empty. So for the whole class of API-level failures — auth, rate limits, model
    access — the fallback string is what surfaces. Confirmed by running the same claude invocation
    manually in the workspace container (exit 0 on a valid key) and reading the failing run's
    transcript, which carries "error":"authentication_failed" / Error: 401 API key is invalid. on
    stdout.

    Suggested fix

    Call persistFailedEngineEvents for any AgentEngineError, not only agent_session_corrupt. The
    rethrow at dispatcher.ts:231 already forwards error.details, so moving the single call into the
    AgentEngineError branch of the outer catch covers both paths without persisting twice.

    The headline message is a separate question — promoting the parsed error result into it would make
    the failure read 401 API key is invalid rather than exited with status 1. Worth doing, but it
    touches the parser rather than the dispatcher, so probably its own change.

    Related

    #328 (fix(github): name the kickstart failure instead of returning a bare 500) is the same class
    of problem on a different path. The argument made there — that the first flow a new installation
    runs is the worst place for an opaque answer — applies here too.

    Happy to open a PR for the dispatcher change if it is wanted.

  3. andreykeycee commented on Sep 15, 2026

    @andreykeycee
    ContributorAuthor

    Root cause and fix in #384

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions