fix(llm): stop passing an httpx.Timeout to the Anthropic SDK - #588
Merged
Merged
Conversation
Every media_editor and reel call in production has failed since
2026-09-18 00:29 with the error "Connection error." The sub-session ran to
completion -- you can watch it finish and write its final report -- and then
the result never reached the parent session. 64 of them in three days; the
task failure rate went from 4.1% to 16.3% and the whole delta is these two
tools.
The sub-session itself was never the problem. session_post ends by asking an
LLM to extract the media URLs out of the finished session, and that is a
structured-output call. anthropic 1.x moved its transport from `httpx` to
`httpx2`, and it now rejects a foreign `httpx.Timeout` in `_build_request`:
TypeError: Invalid `timeout` argument; `httpx.Timeout` is from the
`httpx` package, but this SDK uses `httpx2`.
The SDK's retry loop catches that TypeError, retries twice, and re-raises it
as APIConnectionError -- whose message is just "Connection error.". Hence a
connection-shaped error that failed in 1.4s and was never a connection
problem at all.
Only these two call sites passed a Timeout object, which is why the damage
looked so arbitrary: ordinary streaming chat kept working perfectly while
every Anthropic structured-output call failed. Passing plain seconds is
accepted by both SDK generations, so it cannot break again on the next
transport swap. The 10s connect cap falls back to the SDK default, which is
stricter.
Why now: the Modal image installs with pip_install_from_pyproject, and
pyproject asks for `anthropic>=0.74.0` with no upper bound. Dropping the
unused `runway` dependency (c88917e) changed pyproject, which invalidated
Modal's cached dependency layer and re-resolved every unpinned package for
the first time in months -- anthropic went 0.x -> 1.6.0 in that rebuild.
Note this means requirements.lock does not govern what actually runs: it
pins anthropic==0.74.1 while the deployed container has 1.6.0.
Verified in the real production image (anthropic 1.6.0, httpx2 2.13.0): the
exact session_post extraction call fails with the TypeError before this
change and returns parsed outputs after it.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The outage
Every
media_editorandreelcall in production has failed since 2026-09-18 00:29 with the errorConnection error.The sub-session runs to completion — you can open it and watch it finish and write its final report — and then the result never reaches the parent session.llm_callsareclaude-sonnet-5carryingMEDIA_SYSTEM_PROMPT; zero before 9/18What was actually broken
The sub-session was never the problem.
session_postfinishes by asking an LLM to extract the media URLs out of the completed session — and that is a structured-output call.anthropic 1.x moved its transport from
httpxtohttpx2, and now rejects a foreignhttpx.Timeoutin_build_request:The SDK's retry loop catches that TypeError, retries twice, and re-raises it as
APIConnectionError— whose message is just"Connection error.". Hence a connection-shaped error that failed in 1.4 seconds and was never a connection problem at all.Only two call sites in the whole repo passed a Timeout object, both in
AnthropicProvider, both on the structured-output path. That is why the damage looked so arbitrary: ordinary streaming chat kept working perfectly while every Anthropic structured-output call failed.In session
6aaecd17, child session6ab0410dfinished cleanly at 20:51:49 with the final master rendered; the extraction call failed at 20:51:54, five seconds later.Why it started on 9/18
The Modal image installs with
pip_install_from_pyproject, andpyproject.tomlasks foranthropic>=0.74.0with no upper bound. Commit c88917e dropped the unusedrunwaydependency — that changedpyproject.toml, which invalidated Modal's cached dependency layer and re-resolved every unpinned package for the first time in months.anthropicwent0.x → 1.6.0in that rebuild.Worth noting separately:
requirements.lockdoes not govern what actually runs. It pinsanthropic==0.74.1while the deployed container has1.6.0. CI installs from the lockfile; the Modal image does not.The fix
Pass plain seconds instead of a Timeout object. Accepted by both SDK generations, so it cannot break again on the next transport swap. The 10s connect cap falls back to the SDK default, which is stricter.
Verification
Run inside the real production image (anthropic 1.6.0, httpx2 2.13.0):
session_postextraction call fails with theTypeErroraboveThe added test fails if anyone reintroduces a constructed Timeout — confirmed with a negative control.
🤖 Generated with Claude Code