Skip to content
This repository was archived by the owner on Sep 22, 2026. It is now read-only.
This repository was archived by the owner on Sep 22, 2026. It is now read-only.

Trim governor goes blind after a failed call, so an over-limit conversation can wedge permanently #1477

Description

@njbrake

Note: this issue was drafted by Claude via back-and-forth with @njbrake. The reasoning and decisions are his; the prose is Claude's.

Summary

The trim governor's prompt-size signal is only refreshed on a successful LLM response. Once calls start failing, the signal freezes at the last good value, so the proactive trim in process_message cannot see that the prompt has grown past the limit. A conversation that genuinely exceeds the model's context window can therefore wedge permanently: every message fails, no trim ever fires, and nothing recovers it except the 216-user-turn compaction backstop happening to come due.

This is the same failure mode #1456 described one layer down. That PR fixed the governor reading an undercounted number; this is the governor reading a stale number.

Mechanism

core.py records the prompt size only inside the success path:

if response.usage and response.usage.input_tokens:
    prompt_tokens = response.usage.input_tokens + cache_create + cache_read
    self._last_input_tokens = prompt_tokens
    _remember_input_tokens(self.user.id, prompt_tokens)

A fresh ClawboltAgent is built per message, so the next turn's proactive trim reads _recall_input_tokens(user_id) from the process-local LRU. If every call since the last success failed, that value is however old the last success is.

Two things make this worse than "one turn behind":

  1. _tokens_for is a no-op on the un-trimmed list. It scales the recorded count by a content-length ratio, but orig_len is computed from the same messages list being measured, so _tokens_for(messages) returns actual_input_tokens unchanged. The check is effectively "was the previous prompt over the trigger", not "is this one". A single turn that adds a large tool result (a gmail_list_recent, an appfolio_list_work_orders, several analyze_photo results) can jump straight past the trigger without the governor ever seeing a number above it.
  2. Failure freezes the only correcting signal. Because failures write nothing, the loop cannot self-correct. Prompt over limit means every call fails, which means no usage row, which means the stale value persists, which means no trim.

The reactive path that should catch this (except ContextLengthExceededError then trim and retry) is not a reliable backstop either:

So on the current production path there is a real state a conversation can enter and not leave on its own.

Not the cause of the 2026-08-03 outage

Worth stating so nobody conflates them. That outage was an out-of-credit Anthropic account behind the gateway alias. Jesse's prompt was ~113k tokens, comfortably under the limit, and the proactive trim correctly did not fire. This issue is the adjacent latent bug found while ruling that out.

Options

Not proposing a specific fix; the tradeoffs are worth a decision.

  • Count locally instead of trusting the last response. Estimate the prompt size from the serialized payload actually about to be sent (messages plus tool schemas plus system), so the governor measures this request. Removes the staleness and the one-turn lag together. Costs a serialization pass, though llm_payload_capture already does one on the same path.
  • Update the signal on failure too. Where a failure carries usable usage or a token count in its body, record it. Narrow, and gateways that strip bodies defeat it.
  • Make the trim trigger on a size the governor can always compute. Fall back to the chars/4-plus-overhead heuristic whenever the recorded value is older than the current message count implies, rather than trusting a stale exact number over a fresh estimate.
  • Bound the staleness. Store the message seq alongside the token count and treat the value as unusable once the conversation has advanced more than N turns past it.
  • Belt and braces on the reactive path. Make the trim-and-retry fire on any unclassified provider 400 when the prompt is above some floor, rather than only on a recognized overflow message. Riskier: it retries genuinely malformed requests.

Pointers

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workinglayer: agentAgent loop, tools, orchestration

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions