Skip to content

design(core): reactorRetry retries HTTP 4xx refusals such as an expired delegation, so a query fails only after 16 requests and about 20 seconds #646

Description

@b3hr4d

DESIGN (not implemented: the fix changes behaviour the maintainer should choose).

What happens

isRetryableReactorError treats every agent error without a reject code as retryable ("transport, protocol and certificate failures"). An HTTP error from the replica or boundary node surfaces as a ProtocolError with an HttpErrorCode (status, statusText, bodyText), so a 4xx is retried like a 5xx. HttpAgent already retries a non-2xx answer retryTimes (3) times with its own backoff before it throws, and reactorRetry (the default query retry of defineReactor's QueryClient) then repeats the whole attempt 3 more times.

Measured (a fetch stub answering every query with the status below, retry: reactorRetry, browser mode, fake timers):

response requests sent time until the error surfaces
400 Invalid delegation expiry: … 16 ~22 s
403 16 ~23 s
429 16 ~23 s
500 / 503 16 ~20–21 s

The 429 and 5xx rows are what retrying is for. The 400 and 403 rows are refusals that cannot change: an expired or invalid delegation, a bad signature, a malformed request. The most common real case is a session that has lapsed, where every mounted query spins for about 20 seconds before it shows an error.

Related, not the same: #622 covers update methods run through query hooks, where any retry is a second call. This issue is about refusals that no retry can change, for query methods too.

Why this is a design question

The docs (error-handling.mdx, packages/core.mdx) state the current rule: protocol failures are retryable, and "an unfamiliar transport-level fault is never silently made fatal". Changing it changes documented behaviour.

Options

  1. Do not retry HTTP 4xx except 408 and 429. In isRetryableReactorError, read cause.code.status when cause.code is an HttpErrorCode-shaped object (a status number and name === "HttpErrorCode", not instanceof, so it holds across SDK copies). 5xx, 408, 429, transport and certificate failures stay retryable. Update the two docs pages.
  2. Keep retrying, but fewer times for 4xx (for example once). This halves the wait but still spends it on a refusal.
  3. Keep the rule and document it, suggesting retryTimes: 0 on the agent for apps that prefer fast failures.

Recommendation

Option 1. It keeps the rule's intent (retry only what could change) and only removes cases the replica will repeat. The SDK's own 3 retries still cover a flaky intermediary.

Acceptance

  • isRetryableReactorError returns false for a CallError whose cause carries an HTTP 400/401/403/404 and true for 408, 429 and 5xx.
  • With reactorRetry, a 400 fails after the agent's own attempts (4 requests), not 16.
  • The retry sections of error-handling.mdx and packages/core.mdx describe the new rule.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions