Conversation
joshklop
force-pushed
the
josh/kona-l1-signer-reloader
branch
from
October 5, 2026 20:12
dcae0fc to
d6de3c3
Compare
joshklop
force-pushed
the
josh/kona-l1-signer-reloader
branch
from
October 5, 2026 22:40
13bec91 to
9a3e40a
Compare
Replace the per-call `rpc_timeout` wrappers, the timeout inside `record_call_time`, and the outer timeout around the reset forkchoice update with one deadline on each of the engine client's RPC clients: a tower `MapFutureLayer` on the alloy transport that times out each request. Every request the engine client sends is now bounded, including ones that were not wrapped, and a new call site cannot miss the deadline. The deadline wraps alloy's transport future rather than the hyper client below it: hyper's future completes when the response headers arrive, and alloy reads the body afterwards. The new test covers both a request that is never answered and a response whose body stalls after its headers. The deadline is still per request, so a long reset traversal is not cut off, and a timeout is still a transport error, so the engine retries it as before. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
joshklop
force-pushed
the
josh/kona-l1-signer-reloader
branch
from
October 5, 2026 22:53
9a3e40a to
de042a8
Compare
The L1 watcher sent each signer it read to the network actor over a bounded channel, and the network actor relayed it into a watch channel that the gossip crate created for block validation. Each L1 watcher chain now owns that watch channel, seeded with the signer read at startup, and gossip validation subscribes to it directly. The network actor's relay arm and its closed-channel error are removed. Publishing a signer can no longer block, so the refresh loop no longer reserves channel capacity and re-checks the head before sending, and subscribers are only notified when the signer changes. The gossip builder accepts an external signer receiver. Callers without an L1 watcher, such as `kona-node net` and the gossip example, still pass a fixed address. BREAKING CHANGE: affects code built on the `kona-gossip` and `kona-node-service` crates. `GossipDriverBuilder::build` returns only the driver, `GossipDriverBuilder::with_unsafe_block_signer_receiver` takes a `watch::Receiver<Address>` instead of an `Address`, and `GossipDriverBuilder::new` and `GossipDriver::builder` are no longer `const`. `NetworkHandler` and `NetworkDriver` lose `unsafe_block_signer_sender`, `NetworkActor::new` loses its signer-update receiver, and `L1WatcherChain::new` takes a `watch::Sender<Address>`. To follow a changing signer, create a watch channel, give its sender to the L1 watcher chain, and pass its receiver to `NetworkBuilder::with_unsafe_block_signer`. To use a fixed signer, keep passing an `Address` to `NetworkBuilder::new`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signing moves out of the network actor into a new `SignerActor` between the
sequencer and the network actor. It signs payloads in order and retries
transient signer failures (remote signer transport errors and per-attempt
timeouts) with capped backoff until they succeed, so no payload is skipped or
reordered. A long outage fills the sequencer's 256-payload queue and halts
block production. Other signing errors are fatal, as before. The expected
signer is read from the L1 watcher's watch channel on each attempt. If the
signer rotates between signing and publishing, the network actor publishes
the payload anyway and peers reject it.
This replaces the signing task inside the network actor, which gave each
payload three 2s attempts and then skipped it, leaving a gap in gossip.
The block signer is started only in sequencer mode, before the network binds
its ports. A sequencer without a signer drops its payloads with a warning, as
before. A validator with a signer configured logs that it is ignored, and no
longer health-checks a remote signer at startup.
BREAKING CHANGE: affects code built on the `kona-node-service` crate.
`NetworkBuilder::new` loses its `signer` parameter and
`NetworkBuilder::with_signer` is removed. `NetworkDriver` and
`NetworkHandler` lose `signer`. `NetworkDriverError::BlockSignerStartError`
and `NetworkActorError::{FailedToSignPayload, SigningTask,
MissingUnsafeBlockSigner}` are removed. `NetworkActor::new` takes a receiver
of `SignedPayload` instead of unsigned payloads. To gossip a sequencer's
blocks, start the `BlockSigner` and run a `SignerActor` between the
sequencer's payload channel and the network actor's publish channel.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The `SignerActor` now always holds a signer. A sequencer without one has nothing to sign its blocks with, and its 256-payload queue would never be read, so a node in sequencer mode now fails at startup when no block signer is configured. Before, it started and dropped every payload with a warning, so its blocks were never gossiped. The kona-node sequencer docs and the CLI reference now list the signer flags, and every sequencer example sets one. BREAKING CHANGE: a kona-node started with `--mode sequencer` and none of `--p2p.sequencer.key`, `--p2p.sequencer.key.path`, or `--p2p.signer.endpoint` with `--p2p.signer.address` (or their `KONA_NODE_P2P_*` environment variables) now fails to start. Configure one of them. op-node still allows a sequencer without a signer. For code built on `kona-node-service`, `SignerActor::new` takes a `BlockSignerHandler` instead of an `Option`. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
joshklop
marked this pull request as ready for review
October 5, 2026 23:39
During a long signer outage the gossip queue fills. The sequencer sealed the next block, committed it to op-conductor, and then waited for queue space inside its step, so it stopped answering admin RPC, including op-conductor's `StopSequencer`. A remote signer is typically shared by every sequencer in a cluster, so failing over does not help, but the sequencer still has to answer the conductor. The sequencer now checks for queue space before it seals, and while the queue is full it builds nothing and keeps answering admin queries. It is the queue's only sender, so once there is room the hand-off after sealing cannot wait. No block is skipped: building resumes on the next tick after the queue drains, and the backlog is gossiped in order. The pause and resume are each logged once. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Log every build tick at info, including ticks skipped because the gossip queue is full, instead of tracking the pause in a field to log only its start and end. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A tick skipped because the gossip queue is full means block production is paused, so log it at warn rather than info. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move retries of transient signing failures from the SignerActor into the remote signer, so the actor only sees errors that retrying cannot fix. The remote signer's HTTP client now has a request timeout, which also bounds the startup health check and clients rebuilt for new certificates. Its signing request is retried with exponential backoff until it succeeds when the signer is unreachable, does not answer in time, or answers with an error. A request or response that cannot be encoded or decoded is not retried. The retry is in the signing call rather than in transport middleware because the same client makes the startup health check, which should fail fast. The signer address is now checked once per payload instead of once per attempt. If the unsafe block signer rotates while a request is being retried, the signature returned is for the retired address and the next payload fails with InvalidAddress, as a rotation already did. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #23125 as a proposed simplification of its engine RPC deadlines. Draft while we iterate.
#23125 bounds engine RPCs in three separate places:
rpc_timeoutaround individual call sitesrecord_call_timeA new call site without the wrapper has no deadline. This PR removes all three and puts one deadline on each of the engine client's two RPC clients, the JWT Engine API client and its L1 client. Both get the same alloy transport layer, a stock tower
MapFutureLayerthat wraps each request in a timeout. The call sites are back to their form ondevelop.Every request from the engine client now has a deadline, including ones #23125 does not wrap:
eth_getProofforoptimism_outputAtBlockThe deadline is on alloy's transport future, not on the hyper client below it. hyper's future completes when the response headers arrive, and alloy reads the body after that. A timeout at the hyper level would let an endpoint that stalls mid-body hang the call. The new test covers both a request that is never answered and a response whose body stalls after its headers.
Behavior otherwise matches #23125:
new_payload_v1keeps therecord_call_timewrapper that kona-node: recover from RPC failures without cascading across chains #23125 added, which now only records its latency metric like the other Engine API methods.The L1 watcher's timeouts are unchanged. They use the node's shared L1 provider, which derivation also uses, and a deadline there needs a separate decision.
It also proposes restructuring how the sequencer's blocks are signed:
SignerActorbetween the sequencer and the network actor. kona-node: recover from RPC failures without cascading across chains #23125 signs in a task inside the network actor that gives each payload three 2s attempts and then skips it, so a transient signer outage leaves a gap in gossip. TheSignerActorsigns payloads in order and retries transient failures (remote signer transport errors and per-attempt timeouts) with backoff until they succeed, so nothing is skipped. Other signing errors are fatal, as before. If the signer rotates between signing and publishing, the payload is still published and peers reject it. The network actor goes back to only publishing.Effect on operators:
--p2p.sequencer.key,--p2p.sequencer.key.path, or--p2p.signer.endpointwith--p2p.signer.address). Before, it started and dropped every payload, so its blocks were never gossiped. op-node still allows this, but there it is only useful with p2p disabled, which kona does not support. The kona-node sequencer docs and CLI reference now list the signer flags.BREAKING CHANGE: a kona-node started with
--mode sequencerand no block signer now fails to start; configure one of the signer flags above. The rest affects code built on thekona-gossipandkona-node-servicecrates.GossipDriverBuilder::buildreturns only the driver,GossipDriverBuilder::with_unsafe_block_signer_receivertakes awatch::Receiver<Address>, andGossipDriverBuilder::newandGossipDriver::builderare no longerconst.NetworkBuilder::newloses itssignerparameter, andNetworkBuilder::with_signeris removed.NetworkDriverandNetworkHandlerlosesignerandunsafe_block_signer_sender.NetworkActor::newloses its signer-update receiver and takes a receiver ofSignedPayloadinstead of unsigned payloads.L1WatcherChain::newtakes awatch::Sender<Address>.NetworkDriverError::BlockSignerStartErrorandNetworkActorError::{FailedToSignPayload, SigningTask, MissingUnsafeBlockSigner}are removed.SignerActor::newtakes aBlockSignerHandler; the actor always has a signer.UnsafePayloadGossipClienthas a new required method,has_capacity, which the sequencer checks before sealing.To migrate: give the L1 watcher chain the sender of a signer watch channel and pass its receiver to
NetworkBuilder::with_unsafe_block_signer. To gossip a sequencer's blocks, start theBlockSignerand run aSignerActorbetween the sequencer's payload channel and the network actor's publish channel. Callers with a fixed signer keep passing anAddresstoNetworkBuilder::new.Reviews:
rust-code-reviewerwas run on each commit. For the transport commit, its main finding, that a hyper-level timeout misses a stalled response body, is fixed.deletion-reviewerwas run on the signer commits. It found no surviving references to the removed APIs outside code, and its findings are applied: the breaking changes are marked, a dead error variant is removed, and the kona-node design doc lists the signer actor.🤖 Generated with Claude Code