Skip to content

Wait out transient provider failures instead of aborting the run - #59

Closed
TON14 wants to merge 3 commits into
AMAP-ML:mainfrom
TON14:feat/provider-retry
Closed

TON14 wants to merge 3 commits into
AMAP-ML:mainfrom
TON14:feat/provider-retry

Conversation

@TON14

@TON14 TON14 commented Aug 20, 2026

Copy link
Copy Markdown
Contributor

What this does

Waits out transient provider failures — 429/overload and network drops — with exponential backoff instead of aborting the run. On a long-horizon task, the old behaviour threw away hours of completed rounds for a condition that clears on its own.

Extracted from #57, where it was developed during live Windows testing but has nothing Windows-specific in it.

Behaviour

  • _run_role_episode retries only the two genuinely transient failure kinds, rate_limit and network: 60s doubling to a 900s cap, up to 8 attempts or 2 hours total.
  • Terminal kinds — authentication, quota, model_unavailable, timeout — still fail fast; retrying those wastes time or masks a real hang.
  • The wait is visible, not silent: an agent_runtime_retry event and a role_retry progress record carry the attempt number, the delay, and the provider's own message, so the dashboard shows "waiting, attempt 3" rather than nothing. A stop during backoff still returns a cancelled episode.
  • Tunable via LH_HARNESS_PROVIDER_RETRY_{MAX_ATTEMPTS,BASE_SECONDS,CAP_SECONDS,MAX_TOTAL_SECONDS}; MAX_ATTEMPTS=0 restores the old fail-fast behaviour exactly.

Classification also learns three subscription-limit shapes that previously read as success (the episode "completed" with an empty result): HTTP 529; a result whose is_error is false but whose api_error_status is 429; and a rate_limit_event record with status "rejected".

Testing

  • Full suite passes on this branch (404 passed, 2 skipped on Linux).
  • Verified live inside Windows support: run the harness natively on Windows #57's long-horizon runs on both Linux and Windows, where real 429s from the provider were waited out and the runs completed.
  • Honest gap: the backoff loop itself has no dedicated unit tests yet — the suite covers the classification module but not the retry timing/event emission. Happy to add them in this PR if wanted.

A run died the moment the provider answered 429, overloaded, or dropped the
connection, which on a long-horizon task throws away hours of work for a
condition that clears on its own.

`_run_role_episode` now backs off exponentially for the two failure kinds that
are actually transient -- rate_limit and network -- and retries: 60s doubling to
a 900s cap, up to 8 attempts or 2 hours total, tunable through
LH_HARNESS_PROVIDER_RETRY_{MAX_ATTEMPTS,BASE_SECONDS,CAP_SECONDS,MAX_TOTAL_SECONDS}
with MAX_ATTEMPTS=0 restoring the old fail-fast behaviour. Terminal kinds
(authentication, quota, model_unavailable, timeout) return immediately as
before; retrying those wastes time or masks a real hang. The wait is visible:
an `agent_runtime_retry` event and a `role_retry` progress record carry the
attempt, the delay and the provider's own message, and a cancel during backoff
still returns a cancelled episode.

Classification also learns the subscription-limit shapes that were previously
read as success: 529, a session/usage limit result whose `is_error` is false but
whose `api_error_status` is 429, and a `rate_limit_event` record whose status is
"rejected".
TON14 added 2 commits August 20, 2026 20:51
The PR's own history is the argument: the branch was verified green on each
platform by hand, and each round of hand-verification still found something
the other platform could not see (a POSIX-only test guard, a cmd.exe-only
command-line limit). A matrix of ubuntu + windows at both ends of
requires-python (3.10 and 3.14) makes that check automatic for every push and
pull request.

The suite needs no Node toolchain -- the Web bundle is a packaging artifact --
so the job is checkout, setup-python, `pip install -e ".[test]"`, pytest. The
Windows symlink fixtures skip themselves on runners without
SeCreateSymbolicLinkPrivilege, which is expected and green.
…ands

The windows-latest lanes exercise platform support this branch does not
carry: it is based on a main whose supervisor still calls os.killpg and
whose agent stubs are #!/bin/sh scripts, so those lanes fail on known
pre-existing breakage rather than on anything in this change. AMAP-ML#57 brings
the Windows support together with the full two-platform matrix; when it
merges, its version of this workflow supersedes this one.
@TON14

TON14 commented Sep 17, 2026

Copy link
Copy Markdown
Contributor Author

This PR was auto-closed when I cleaned up the head branches in my fork. Upstream has been quiet since 2026-08-20, so I keep maintaining this work (together with the rest of my changes) in my fork's main: https://github.com/TON14/LongHorizon-Harness. Happy to rebase and reopen if upstream activity resumes — thanks for the project!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant