Skip to content

logs source: ask again for a block the upstream cannot serve yet - #399

Merged
l0gun0v merged 3 commits into
mainfrom
logs-source-retry-not-ready
Oct 5, 2026
Merged

l0gun0v merged 3 commits into
mainfrom
logs-source-retry-not-ready

Conversation

@l0gun0v

@l0gun0v l0gun0v commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

The local logs subscription source asks again for a block when the upstreams at its height report it as not ready yet, instead of skipping it.

Seen on ethereum drpc-core US-Central: the only upstream with eth_getLogs there is erigon 3. Erigon dispatches newHeads before it commits the block, and its eth_getLogs reads only committed data, so for the 0.7–1 s commit window it answers block range extends beyond current head block / requested block range [N,N] is beyond latest executed block. The source asks within milliseconds, so ~100% of blocks were skipped there. Gnosis, sepolia and hoodi on erigon behave the same, and cosmos-evm chains answer block not found for hash while they index a fresh block.

Spec

  • Not-ready answers. block range extends beyond current head block, beyond latest executed block, block not found, unknown block, header not found, could not find results for height (case-insensitive substring).
  • Retry. Such an answer moves on to the next upstream, as any error does, and does not use up one of the 3 attempts. When no upstream at the height could serve the block and at least one said it is not ready, the block is asked again from the top of the rating list with a backoff of 100 ms doubling to 1 s.
  • Deadline. One expected-block-time, clamped to 1–3 s, counted from the head's arrival (BlockUpdate.Seen), so blocks queued behind a waiting one do not add up. After it the block is skipped as before. Other errors are not waited on.
  • Observability. Skips after not-ready answers get reason="not_ready". Skip warnings carry the last upstream and its error (upstream, upstream_error): the selection error alone hid the error of the upstream that was asked. New counter nodecore_logs_source_not_ready_retries_total{chain}.
  • Compatibility. No config changes. Only the local logs source is affected.

Changes

File Change
internal/upstreams/flow/logs_source.go blockNotReady, logsNotReadyWait; fetchBlockLogs retries not-ready rounds until the deadline; skip reason and log fields; new metric
internal/upstreams/flow/subengine/blockupdates.go BlockUpdate.Seen (head arrival time)
internal/upstreams/flow/logs_source_internal_test.go error matching; retry until served; deadline; fall-through to an upstream that can serve; other errors not delayed; end-to-end retry
docs/nodecore/08-prometheus-metrics.md, docs/nodecore/13-subscriptions.md not_ready reason, log fields, new metric, behaviour

An upstream may announce a block before it can serve its logs: erigon 3
dispatches newHeads from the execution overlay before it commits the block,
and its eth_getLogs, which reads only committed data, answers "block range
extends beyond current head block" / "block not found" for the 0.7-1 s commit
window. The logs source asked within milliseconds, so on a host where erigon
is the only eth_getLogs upstream (cherry-us-bcn-05, gnosis on fornex) almost
every block was skipped. Cosmos-evm nodes answer "block not found for hash"
the same way while they index a fresh block.

- Such an answer moves on to the next upstream, as any error does. When no
  upstream at the height could serve the block, they are asked again with a
  backoff (100 ms .. 1 s) until one block time after the head arrived,
  clamped to 1..3 s; other errors are not waited on. BlockUpdate carries the
  head's arrival time for the deadline.
- Skips after such answers are counted as reason "not_ready"; the skip warning
  logs the last upstream and its error instead of only the selection error.
- Metric logs_source_not_ready_retries_total.
@l0gun0v
l0gun0v marked this pull request as ready for review October 5, 2026 09:43
Comment thread internal/upstreams/flow/logs_source.go
denis and others added 2 commits October 5, 2026 13:36
Each not-ready round built a fresh strategy, so an upstream at the same
height that answered another error (5xx, rate limit) was asked again every
round and used up an attempt each time: with erigon not ready and such an
upstream next to it, the block was skipped as upstream_error after ~300 ms,
before erigon's commit window ended. Such upstreams are now left out of
later rounds, so each of them counts once, as on a single walk down the list.
@l0gun0v
l0gun0v requested a review from KirillPamPam October 5, 2026 10:47
@l0gun0v
l0gun0v merged commit 877611b into main Oct 5, 2026
5 checks passed
@l0gun0v
l0gun0v deleted the logs-source-retry-not-ready branch October 5, 2026 10:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants