Skip to content

Fix deploys failing on PyPI timeouts: install from the lockfile, drop unused runway dep - #585

Merged
genekogan merged 13 commits into
stagingfrom
fix/ci-lockfile-flake-20260917
Sep 17, 2026
Merged

genekogan merged 13 commits into
stagingfrom
fix/ci-lockfile-flake-20260917

Conversation

@genekogan

Copy link
Copy Markdown
Contributor

The failure

Both jobs on main@d413cf80 died in rye sync, ~2 minutes in:

Deploy to Modal Production
  error: Failed to fetch: `https://pypi.org/simple/farcaster/`
Tool Pricing Drift Check
  error: Failed to fetch: `https://pypi.org/simple/runway/`
  Caused by: Request failed after 3 retries -> operation timed out

Two different packages, two jobs, the same minute. That is PyPI/runner flakiness during dependency resolution, not a broken dependency. main has not deployed to Modal since 2026-09-16, so d413cf80 is currently undeployed in production.

Why CI was resolving at all

Every job ran bare rye sync, which regenerates requirements.lock from scratch before installing — re-resolving ~250 packages over the network on each deploy. The lockfile is committed specifically so that does not have to happen.

The fix

  1. rye sync --no-lock in all six workflows. Installs straight from the committed lockfile and makes no resolution requests, so deploys are deterministic and cannot fail this way.
  2. New Lockfile Drift Check workflow, covering the one thing --no-lock gives up: a pyproject.toml edit without a re-lock would otherwise install the old dependency set silently. It re-resolves and fails if the committed locks are stale. Advisory (continue-on-error) on purpose — it is now the only job that talks to PyPI, so it must never block a deploy. Also raises UV_HTTP_TIMEOUT from uv's 30s default to 180s.
  3. Drops the runway dependency. It is the Onica/AWS CloudFormation tool, added in 6d14118 in the same diff that bumped runwayml — a rye add runway that meant runwayml. Nothing imports it. It pulled 21 packages of CloudFormation tooling (cfn-lint, troposphere, aws-sam-translator, awacs, docker, pyopenssl, yamllint, …) into the production image.

Verification

  • rye lock on main reproduces the committed requirements.lock and requirements-dev.lock byte-for-byte, so --no-lock installs exactly what the old path resolved to.
  • No direct import of runway or of any dropped transitive package anywhere in eve/, tests/ or scripts/.
  • runwayml==3.11.0 untouched; the runway removal is a pure removal — not one surviving package changed version.
  • requirements.lock: 252 → 231 packages.

Merging this deploys d413cf80 along with it.

🤖 Generated with Claude Code

genekogan and others added 13 commits August 23, 2026 14:52
…sk doc

refund_manna() deduped with an upsert on {task, type:"refund"}, race-proof only
"because of the UNIQUE index on {task, type} (see
Transaction.ensure_unique_refund_index)".

That function does not exist anywhere in the repo. The index it names DOES
exist in eden-prod (uniq_refund_per_task: {task:1,type:1}, unique, partial on
{type:"refund"}) — but it was applied out of band and is referenced nowhere in
version control. So the dedup is safe today, and would silently stop being safe
on any rebuilt or restored cluster, with no test or startup check to catch it.

That matters because two refund callers race by construction on every cancel of
a running task, in two different processes: Tool.handle_cancel in the API
container and _task_handler's finally block in the Modal worker.

Two changes:

- The claim now happens on the task document via a conditional
  find_one_and_update on _id, so it depends only on the primary key index and
  cannot be undone by missing out-of-band state. The transaction row stays as
  the ledger (still an upsert, so a retry cannot leave two rows), and the claim
  is released if the credit fails, so a retriable error cannot silently eat
  manna the user is owed.

- Transaction.ensure_unique_refund_index now exists and is registered in
  add_cost_indexes.py, specified to match the deployed index exactly so it is a
  no-op against prod. It must stay partial: session LLM charges write many
  type="spend" rows sharing task=<session id>, so a non-partial unique index on
  {task, type} would fail to build.

Tests: the two double-credit cases fail against the previous implementation.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
fix(manna): declare the refund index in code; claim refunds on the task doc
…ollected

The 2-manna floor is charged once per RUN (`_run_floor_charged`), so after the
first turn nothing re-asserts the caller's balance. Cost-plus top-ups in
_persist_assistant_message are deliberately best-effort — a post-hoc charge
failure must not kill a response the user already received — but the failure
was only logged and the loop carried on.

An account funded with ~2 manna therefore bought a full
MAX_PROMPT_SESSION_TURNS (25) tool loop: the floor drains the balance on turn
1, every later top-up raises on insufficient manna and is swallowed, and tool
calls that fail their own check_manna return error results, which is exactly
what keeps the LLM looping. On a large context that is several dollars of
provider spend against 2 manna collected.

This is a regression introduced by the streaming billing fix (#572) that moved
the floor from per-LLM-call to once-per-run; before it, an unfunded account hit
the pre-turn APIError path and the run ended.

The failure is now recorded on the runtime and enforced at the next turn
boundary. That keeps the already-delivered response intact — the property the
best-effort design exists to protect — while refusing to start another paid
turn.

Tests: the stop case fails against the current implementation (the run
continues to max_turns). Two control cases assert the guard does not fire on a
healthy run and is inert when FF_MANNA_BILLING is unset, so it cannot end runs
prematurely.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
fix(billing): stop a run once its cost-plus top-up can no longer be collected
…260916

Retire unused legacy Google Calendar integration
Added two new ObjectId entries to the list.
…ery run

Every CI job ran `rye sync`, which regenerates requirements.lock from
scratch before installing. That made each deploy re-resolve ~250 packages
against PyPI over the network, and it is why deploys just started failing:

  Deploy to Modal Production  -> timed out fetching .../simple/farcaster/
  Tool Pricing Drift Check    -> timed out fetching .../simple/runway/

Two different packages, two jobs, same minute -- so this is PyPI/runner
flakiness during resolution, not a bad dependency. Re-resolving was never
wanted here anyway: requirements.lock is committed precisely so that what
gets deployed is what was pinned.

`rye sync --no-lock` installs straight from requirements.lock and makes no
resolution requests at all, so deploys become both deterministic and immune
to this class of failure.

The tradeoff is that a pyproject.toml edit without a matching re-lock would
now install the old dependency set silently, so Lockfile Drift Check is
added to catch exactly that. It is advisory (continue-on-error): it is the
only job that still talks to PyPI, so it must never block a deploy. It also
raises UV_HTTP_TIMEOUT from uv's 30s default to 180s.

Verified: `rye lock` on main reproduces the committed requirements.lock and
requirements-dev.lock byte-for-byte, so --no-lock installs the same set the
old path was resolving to.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`runway` is the Onica/AWS CloudFormation deployment tool. It entered in
6d14118 ("update some deps, peg instructor"), in the same diff that bumped
`runwayml>=3.7.1` -> `>=3.11.0` -- i.e. it was almost certainly a
`rye add runway` meant for the Runway ML SDK, which was already a
dependency and is the thing the runway/runway2/runway3 tools actually use.

Nothing imports it. It cost 21 packages of unrelated CloudFormation
tooling in the production image: cfn-lint, aws-sam-translator, troposphere,
awacs, docker, pyopenssl, yamllint, python-hcl2 and friends. It was also
one of the two packages whose PyPI metadata fetch timed out in the deploy
failures this fixes, because it widened the resolution surface for no
reason.

Verified: no direct import of runway or of any dropped transitive package;
runwayml==3.11.0 is untouched; and the re-lock is a pure removal -- not one
surviving package changed version.

  requirements.lock: 252 -> 231 packages

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The new Lockfile Drift Check caught this on its first run, and it is the
thing that would have broken the deploy: the committed lockfiles were
generated on macOS with `universal: false`, so they are macOS-shaped. They
pin torch==2.7.0 and contain zero nvidia-* or triton packages, because
those are Linux-only wheels that a macOS resolve never sees.

That was survivable only because CI re-resolved from scratch on every run,
on Linux, and quietly produced a *different* dependency set than the one
committed. Switching the deploys to `rye sync --no-lock` without this
commit would have installed the macOS set on Linux: torch with none of its
CUDA runtime.

`rye lock --universal` resolves for every platform at once and emits
environment markers, e.g.

    triton==3.3.0 ; platform_machine == 'x86_64' and platform_system == 'Linux'

so one committed lockfile is correct on both the runner and a dev laptop.
Rye persists `universal: true` in the lockfile header, so later `rye lock`
runs stay universal without anyone having to remember the flag.

Verified with uv dry-run installs of the full lock:
  x86_64-manylinux_2_28  -> resolves
  aarch64-apple-darwin   -> resolves

(manylinux_2_17 does not, because the nvidia wheels are 2_28+. Both
ubuntu-latest runners and the Modal images are well past that.)

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Lockfile Drift Check failed a second time, on marker spelling rather than
on content: rye 0.42 (local) emits

    nvidia-cublas-cu12==12.6.4.1 ; platform_machine == 'x86_64' and platform_system == 'Linux'

where the newer uv inside rye 0.44 (what setup-rye@v4 installs as "latest")
normalizes the same constraint to

    nvidia-cublas-cu12==12.6.4.1 ; platform_machine == 'x86_64' and sys_platform == 'linux'

and likewise collapses grpclib's python_full_version split and pywin32's
redundant Windows marker. Semantically identical, textually different --
and a byte-comparing drift check cannot tell those apart, so it would have
failed on every version skew between a laptop and the runner.

So the version is now part of the contract: setup-rye is pinned to 0.44.0
in all seven workflows, and the lockfiles are regenerated with that exact
version. 0.44.0 is rye's final release (the project is in maintenance,
superseded by uv), so this pin is stable rather than a snapshot that will
drift.

Lockfile Drift Check also now runs `rye lock --universal` to match how the
committed locks are generated; without the flag it would re-resolve
non-universally and report the entire nvidia/triton block as drift.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Third and last cause of Lockfile Drift Check disagreeing with a laptop: the
repo pins no interpreter. pyproject.toml declares no `requires-python` and
there was no .python-version, so rye picks a default and the universal
resolve takes that interpreter as its floor. CI landed on cpython@3.12.9
and a laptop on cpython@3.13.2, which is enough to change which packages
are recorded as needing typing-extensions (aiosignal, anyio, referencing
and starlette all require it conditionally below 3.13).

Pinned to 3.12.9, which is what CI was already choosing, so this changes
nothing about how CI resolves -- it just stops a laptop from disagreeing
with it. The locks are regenerated against it and now match CI exactly.

Worth knowing for context: the Modal image does NOT install from these
lockfiles. eve/api/api.py builds with
`.pip_install_from_pyproject(pyproject.toml)` on python_version="3.11", so
it re-resolves at image-build time. requirements.lock governs the CI
runner's own venv -- the one that runs `modal deploy` and
`eve tool price-check` -- not production. Dropping `runway` does still
slim the deployed image, because that comes from pyproject.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@genekogan genekogan mentioned this pull request Sep 17, 2026
@genekogan
genekogan changed the base branch from main to staging September 17, 2026 21:54
@genekogan

Copy link
Copy Markdown
Contributor Author

Retargeted from main to staging at Gene's request, so this goes through the normal staging-first flow. The branch already contains all of main, so this one merge both catches staging up (it was 8 behind) and lands the CI fix — which matters, because a staging deploy without the fix would hit the same PyPI timeout.

Supersedes #586, which is now empty.

🤖 Generated with Claude Code

@genekogan
genekogan merged commit 5fccc48 into staging Sep 17, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant