Conversation
A run died the moment the provider answered 429, overloaded, or dropped the
connection, which on a long-horizon task throws away hours of work for a
condition that clears on its own.
`_run_role_episode` now backs off exponentially for the two failure kinds that
are actually transient -- rate_limit and network -- and retries: 60s doubling to
a 900s cap, up to 8 attempts or 2 hours total, tunable through
LH_HARNESS_PROVIDER_RETRY_{MAX_ATTEMPTS,BASE_SECONDS,CAP_SECONDS,MAX_TOTAL_SECONDS}
with MAX_ATTEMPTS=0 restoring the old fail-fast behaviour. Terminal kinds
(authentication, quota, model_unavailable, timeout) return immediately as
before; retrying those wastes time or masks a real hang. The wait is visible:
an `agent_runtime_retry` event and a `role_retry` progress record carry the
attempt, the delay and the provider's own message, and a cancel during backoff
still returns a cancelled episode.
Classification also learns the subscription-limit shapes that were previously
read as success: 529, a session/usage limit result whose `is_error` is false but
whose `api_error_status` is 429, and a `rate_limit_event` record whose status is
"rejected".
The PR's own history is the argument: the branch was verified green on each platform by hand, and each round of hand-verification still found something the other platform could not see (a POSIX-only test guard, a cmd.exe-only command-line limit). A matrix of ubuntu + windows at both ends of requires-python (3.10 and 3.14) makes that check automatic for every push and pull request. The suite needs no Node toolchain -- the Web bundle is a packaging artifact -- so the job is checkout, setup-python, `pip install -e ".[test]"`, pytest. The Windows symlink fixtures skip themselves on runners without SeCreateSymbolicLinkPrivilege, which is expected and green.
…ands The windows-latest lanes exercise platform support this branch does not carry: it is based on a main whose supervisor still calls os.killpg and whose agent stubs are #!/bin/sh scripts, so those lanes fail on known pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings the Windows support together with the full two-platform matrix; when it merges, its version of this workflow supersedes this one.
Contributor
Author
|
This PR was auto-closed when I cleaned up the head branches in my fork. Upstream has been quiet since 2026-08-20, so I keep maintaining this work (together with the rest of my changes) in my fork's |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this does
Waits out transient provider failures — 429/overload and network drops — with exponential backoff instead of aborting the run. On a long-horizon task, the old behaviour threw away hours of completed rounds for a condition that clears on its own.
Extracted from #57, where it was developed during live Windows testing but has nothing Windows-specific in it.
Behaviour
_run_role_episoderetries only the two genuinely transient failure kinds,rate_limitandnetwork: 60s doubling to a 900s cap, up to 8 attempts or 2 hours total.authentication,quota,model_unavailable,timeout— still fail fast; retrying those wastes time or masks a real hang.agent_runtime_retryevent and arole_retryprogress record carry the attempt number, the delay, and the provider's own message, so the dashboard shows "waiting, attempt 3" rather than nothing. A stop during backoff still returns a cancelled episode.LH_HARNESS_PROVIDER_RETRY_{MAX_ATTEMPTS,BASE_SECONDS,CAP_SECONDS,MAX_TOTAL_SECONDS};MAX_ATTEMPTS=0restores the old fail-fast behaviour exactly.Classification also learns three subscription-limit shapes that previously read as success (the episode "completed" with an empty result): HTTP 529; a result whose
is_erroris false but whoseapi_error_statusis 429; and arate_limit_eventrecord with status"rejected".Testing