logs source: ask again for a block the upstream cannot serve yet - #399
Merged
Merged
Conversation
An upstream may announce a block before it can serve its logs: erigon 3 dispatches newHeads from the execution overlay before it commits the block, and its eth_getLogs, which reads only committed data, answers "block range extends beyond current head block" / "block not found" for the 0.7-1 s commit window. The logs source asked within milliseconds, so on a host where erigon is the only eth_getLogs upstream (cherry-us-bcn-05, gnosis on fornex) almost every block was skipped. Cosmos-evm nodes answer "block not found for hash" the same way while they index a fresh block. - Such an answer moves on to the next upstream, as any error does. When no upstream at the height could serve the block, they are asked again with a backoff (100 ms .. 1 s) until one block time after the head arrived, clamped to 1..3 s; other errors are not waited on. BlockUpdate carries the head's arrival time for the deadline. - Skips after such answers are counted as reason "not_ready"; the skip warning logs the last upstream and its error instead of only the selection error. - Metric logs_source_not_ready_retries_total.
l0gun0v
marked this pull request as ready for review
October 5, 2026 09:43
KirillPamPam
reviewed
Oct 5, 2026
Each not-ready round built a fresh strategy, so an upstream at the same height that answered another error (5xx, rate limit) was asked again every round and used up an attempt each time: with erigon not ready and such an upstream next to it, the block was skipped as upstream_error after ~300 ms, before erigon's commit window ended. Such upstreams are now left out of later rounds, so each of them counts once, as on a single walk down the list.
KirillPamPam
approved these changes
Oct 5, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The local
logssubscription source asks again for a block when the upstreams at its height report it as not ready yet, instead of skipping it.Seen on ethereum
drpc-coreUS-Central: the only upstream witheth_getLogsthere is erigon 3. Erigon dispatchesnewHeadsbefore it commits the block, and itseth_getLogsreads only committed data, so for the 0.7–1 s commit window it answersblock range extends beyond current head block/requested block range [N,N] is beyond latest executed block. The source asks within milliseconds, so ~100% of blocks were skipped there. Gnosis, sepolia and hoodi on erigon behave the same, and cosmos-evm chains answerblock not found for hashwhile they index a fresh block.Spec
block range extends beyond current head block,beyond latest executed block,block not found,unknown block,header not found,could not find results for height(case-insensitive substring).expected-block-time, clamped to 1–3 s, counted from the head's arrival (BlockUpdate.Seen), so blocks queued behind a waiting one do not add up. After it the block is skipped as before. Other errors are not waited on.reason="not_ready". Skip warnings carry the last upstream and its error (upstream,upstream_error): the selection error alone hid the error of the upstream that was asked. New counternodecore_logs_source_not_ready_retries_total{chain}.logssource is affected.Changes
internal/upstreams/flow/logs_source.goblockNotReady,logsNotReadyWait;fetchBlockLogsretries not-ready rounds until the deadline; skip reason and log fields; new metricinternal/upstreams/flow/subengine/blockupdates.goBlockUpdate.Seen(head arrival time)internal/upstreams/flow/logs_source_internal_test.godocs/nodecore/08-prometheus-metrics.md,docs/nodecore/13-subscriptions.mdnot_readyreason, log fields, new metric, behaviour