Skip to content

fix(llm): stop passing an httpx.Timeout to the Anthropic SDK - #588

Merged
genekogan merged 1 commit into
stagingfrom
fix/anthropic-structured-output-timeout
Sep 20, 2026
Merged

genekogan merged 1 commit into
stagingfrom
fix/anthropic-structured-output-timeout

Conversation

@genekogan

Copy link
Copy Markdown
Contributor

The outage

Every media_editor and reel call in production has failed since 2026-09-18 00:29 with the error Connection error. The sub-session runs to completion — you can open it and watch it finish and write its final report — and then the result never reaches the parent session.

  • 64 failed calls in three days, all of them these two tools
  • task failure rate went 4.1% → 16.3%, and the entire delta is this
  • all 64 failed llm_calls are claude-sonnet-5 carrying MEDIA_SYSTEM_PROMPT; zero before 9/18

What was actually broken

The sub-session was never the problem. session_post finishes by asking an LLM to extract the media URLs out of the completed session — and that is a structured-output call.

anthropic 1.x moved its transport from httpx to httpx2, and now rejects a foreign httpx.Timeout in _build_request:

TypeError: Invalid `timeout` argument; `httpx.Timeout` is from the `httpx`
package, but this SDK uses `httpx2`. Use `httpx2.Timeout` instead.

The SDK's retry loop catches that TypeError, retries twice, and re-raises it as APIConnectionError — whose message is just "Connection error.". Hence a connection-shaped error that failed in 1.4 seconds and was never a connection problem at all.

Only two call sites in the whole repo passed a Timeout object, both in AnthropicProvider, both on the structured-output path. That is why the damage looked so arbitrary: ordinary streaming chat kept working perfectly while every Anthropic structured-output call failed.

In session 6aaecd17, child session 6ab0410d finished cleanly at 20:51:49 with the final master rendered; the extraction call failed at 20:51:54, five seconds later.

Why it started on 9/18

The Modal image installs with pip_install_from_pyproject, and pyproject.toml asks for anthropic>=0.74.0 with no upper bound. Commit c88917e dropped the unused runway dependency — that changed pyproject.toml, which invalidated Modal's cached dependency layer and re-resolved every unpinned package for the first time in months. anthropic went 0.x → 1.6.0 in that rebuild.

Worth noting separately: requirements.lock does not govern what actually runs. It pins anthropic==0.74.1 while the deployed container has 1.6.0. CI installs from the lockfile; the Modal image does not.

The fix

Pass plain seconds instead of a Timeout object. Accepted by both SDK generations, so it cannot break again on the next transport swap. The 10s connect cap falls back to the SDK default, which is stricter.

Verification

Run inside the real production image (anthropic 1.6.0, httpx2 2.13.0):

  • before this change, the exact session_post extraction call fails with the TypeError above
  • after it, the same call returns parsed outputs
  • non-structured streaming chat is unaffected either way

The added test fails if anyone reintroduces a constructed Timeout — confirmed with a negative control.

🤖 Generated with Claude Code

Every media_editor and reel call in production has failed since
2026-09-18 00:29 with the error "Connection error." The sub-session ran to
completion -- you can watch it finish and write its final report -- and then
the result never reached the parent session. 64 of them in three days; the
task failure rate went from 4.1% to 16.3% and the whole delta is these two
tools.

The sub-session itself was never the problem. session_post ends by asking an
LLM to extract the media URLs out of the finished session, and that is a
structured-output call. anthropic 1.x moved its transport from `httpx` to
`httpx2`, and it now rejects a foreign `httpx.Timeout` in `_build_request`:

    TypeError: Invalid `timeout` argument; `httpx.Timeout` is from the
    `httpx` package, but this SDK uses `httpx2`.

The SDK's retry loop catches that TypeError, retries twice, and re-raises it
as APIConnectionError -- whose message is just "Connection error.". Hence a
connection-shaped error that failed in 1.4s and was never a connection
problem at all.

Only these two call sites passed a Timeout object, which is why the damage
looked so arbitrary: ordinary streaming chat kept working perfectly while
every Anthropic structured-output call failed. Passing plain seconds is
accepted by both SDK generations, so it cannot break again on the next
transport swap. The 10s connect cap falls back to the SDK default, which is
stricter.

Why now: the Modal image installs with pip_install_from_pyproject, and
pyproject asks for `anthropic>=0.74.0` with no upper bound. Dropping the
unused `runway` dependency (c88917e) changed pyproject, which invalidated
Modal's cached dependency layer and re-resolved every unpinned package for
the first time in months -- anthropic went 0.x -> 1.6.0 in that rebuild.
Note this means requirements.lock does not govern what actually runs: it
pins anthropic==0.74.1 while the deployed container has 1.6.0.

Verified in the real production image (anthropic 1.6.0, httpx2 2.13.0): the
exact session_post extraction call fails with the TypeError before this
change and returns parsed outputs after it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@genekogan
genekogan merged commit b1fed71 into staging Sep 20, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant