Conversation
…ery run Every CI job ran `rye sync`, which regenerates requirements.lock from scratch before installing. That made each deploy re-resolve ~250 packages against PyPI over the network, and it is why deploys just started failing: Deploy to Modal Production -> timed out fetching .../simple/farcaster/ Tool Pricing Drift Check -> timed out fetching .../simple/runway/ Two different packages, two jobs, same minute -- so this is PyPI/runner flakiness during resolution, not a bad dependency. Re-resolving was never wanted here anyway: requirements.lock is committed precisely so that what gets deployed is what was pinned. `rye sync --no-lock` installs straight from requirements.lock and makes no resolution requests at all, so deploys become both deterministic and immune to this class of failure. The tradeoff is that a pyproject.toml edit without a matching re-lock would now install the old dependency set silently, so Lockfile Drift Check is added to catch exactly that. It is advisory (continue-on-error): it is the only job that still talks to PyPI, so it must never block a deploy. It also raises UV_HTTP_TIMEOUT from uv's 30s default to 180s. Verified: `rye lock` on main reproduces the committed requirements.lock and requirements-dev.lock byte-for-byte, so --no-lock installs the same set the old path was resolving to. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
`runway` is the Onica/AWS CloudFormation deployment tool. It entered in 6d14118 ("update some deps, peg instructor"), in the same diff that bumped `runwayml>=3.7.1` -> `>=3.11.0` -- i.e. it was almost certainly a `rye add runway` meant for the Runway ML SDK, which was already a dependency and is the thing the runway/runway2/runway3 tools actually use. Nothing imports it. It cost 21 packages of unrelated CloudFormation tooling in the production image: cfn-lint, aws-sam-translator, troposphere, awacs, docker, pyopenssl, yamllint, python-hcl2 and friends. It was also one of the two packages whose PyPI metadata fetch timed out in the deploy failures this fixes, because it widened the resolution surface for no reason. Verified: no direct import of runway or of any dropped transitive package; runwayml==3.11.0 is untouched; and the re-lock is a pure removal -- not one surviving package changed version. requirements.lock: 252 -> 231 packages Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The new Lockfile Drift Check caught this on its first run, and it is the
thing that would have broken the deploy: the committed lockfiles were
generated on macOS with `universal: false`, so they are macOS-shaped. They
pin torch==2.7.0 and contain zero nvidia-* or triton packages, because
those are Linux-only wheels that a macOS resolve never sees.
That was survivable only because CI re-resolved from scratch on every run,
on Linux, and quietly produced a *different* dependency set than the one
committed. Switching the deploys to `rye sync --no-lock` without this
commit would have installed the macOS set on Linux: torch with none of its
CUDA runtime.
`rye lock --universal` resolves for every platform at once and emits
environment markers, e.g.
triton==3.3.0 ; platform_machine == 'x86_64' and platform_system == 'Linux'
so one committed lockfile is correct on both the runner and a dev laptop.
Rye persists `universal: true` in the lockfile header, so later `rye lock`
runs stay universal without anyone having to remember the flag.
Verified with uv dry-run installs of the full lock:
x86_64-manylinux_2_28 -> resolves
aarch64-apple-darwin -> resolves
(manylinux_2_17 does not, because the nvidia wheels are 2_28+. Both
ubuntu-latest runners and the Modal images are well past that.)
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Lockfile Drift Check failed a second time, on marker spelling rather than
on content: rye 0.42 (local) emits
nvidia-cublas-cu12==12.6.4.1 ; platform_machine == 'x86_64' and platform_system == 'Linux'
where the newer uv inside rye 0.44 (what setup-rye@v4 installs as "latest")
normalizes the same constraint to
nvidia-cublas-cu12==12.6.4.1 ; platform_machine == 'x86_64' and sys_platform == 'linux'
and likewise collapses grpclib's python_full_version split and pywin32's
redundant Windows marker. Semantically identical, textually different --
and a byte-comparing drift check cannot tell those apart, so it would have
failed on every version skew between a laptop and the runner.
So the version is now part of the contract: setup-rye is pinned to 0.44.0
in all seven workflows, and the lockfiles are regenerated with that exact
version. 0.44.0 is rye's final release (the project is in maintenance,
superseded by uv), so this pin is stable rather than a snapshot that will
drift.
Lockfile Drift Check also now runs `rye lock --universal` to match how the
committed locks are generated; without the flag it would re-resolve
non-universally and report the entire nvidia/triton block as drift.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Third and last cause of Lockfile Drift Check disagreeing with a laptop: the repo pins no interpreter. pyproject.toml declares no `requires-python` and there was no .python-version, so rye picks a default and the universal resolve takes that interpreter as its floor. CI landed on cpython@3.12.9 and a laptop on cpython@3.13.2, which is enough to change which packages are recorded as needing typing-extensions (aiosignal, anyio, referencing and starlette all require it conditionally below 3.13). Pinned to 3.12.9, which is what CI was already choosing, so this changes nothing about how CI resolves -- it just stops a laptop from disagreeing with it. The locks are regenerated against it and now match CI exactly. Worth knowing for context: the Modal image does NOT install from these lockfiles. eve/api/api.py builds with `.pip_install_from_pyproject(pyproject.toml)` on python_version="3.11", so it re-resolves at image-build time. requirements.lock governs the CI runner's own venv -- the one that runs `modal deploy` and `eve tool price-check` -- not production. Dropping `runway` does still slim the deployed image, because that comes from pyproject. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fix deploys failing on PyPI timeouts: install from the lockfile, drop unused runway dep
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Promotion of what was just verified on
staging.mainhas not deployed to Modal since 2026-09-16 —d413cf80is currently undeployed in production, because barerye syncre-resolved ~250 packages against PyPI on every run and timed out. This merge restores the deploy and carriesd413cf80with it.What lands
c3fbcca7rye sync --no-lockinstead of re-resolving; adds advisory Lockfile Drift Checkc88917edrunway— the AWS CloudFormation tool, never imported, added by arye add runwaythat meantrunwayml. Removes 21 packages from the image2d2274ac1b50ca860365b832.python-versionto 3.12.9, which CI was already choosingVerified on staging
Deploy to Modal Staging,Tool Pricing Drift Check(PROD and STAGE) andLockfile Drift Checkall pass onstaging@5fccc488. The install step:928ms, against a 2-minute timeout before. The Modal app deployed.
Note the Modal image itself builds from
pyproject.tomlvia.pip_install_from_pyproject, not from the lockfile, so the lock governs the CI runner's venv. Therunwayremoval still slims the deployed image, since that comes frompyproject.toml.🤖 Generated with Claude Code