Skip to content

kona-node!: bound engine RPCs at the transport and add a SignerActor - #23202

Open
joshklop wants to merge 8 commits into
codex/kona-rpc-recoveryfrom
josh/kona-l1-signer-reloader
Open

joshklop wants to merge 8 commits into
codex/kona-rpc-recoveryfrom
josh/kona-l1-signer-reloader

Conversation

@joshklop

@joshklop joshklop commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #23125 as a proposed simplification of its engine RPC deadlines. Draft while we iterate.

#23125 bounds engine RPCs in three separate places:

  • rpc_timeout around individual call sites
  • a timeout inside record_call_time
  • an outer timeout around the reset forkchoice update

A new call site without the wrapper has no deadline. This PR removes all three and puts one deadline on each of the engine client's two RPC clients, the JWT Engine API client and its L1 client. Both get the same alloy transport layer, a stock tower MapFutureLayer that wraps each request in a timeout. The call sites are back to their form on develop.

Every request from the engine client now has a deadline, including ones #23125 does not wrap:

  • eth_getProof for optimism_outputAtBlock
  • payload bodies and client version
  • capability exchange, so a stalled Engine API endpoint now fails the startup JWT check instead of hanging it

The deadline is on alloy's transport future, not on the hyper client below it. hyper's future completes when the response headers arrive, and alloy reads the body after that. A timeout at the hyper level would let an endpoint that stalls mid-body hang the call. The new test covers both a request that is never answered and a response whose body stalls after its headers.

Behavior otherwise matches #23125:

  • The deadline is still per request, so a long reset traversal is not cut off.
  • A timeout is still a transport error, so the engine classifies it as temporary and retries it as before.
  • new_payload_v1 keeps the record_call_time wrapper that kona-node: recover from RPC failures without cascading across chains #23125 added, which now only records its latency metric like the other Engine API methods.

The L1 watcher's timeouts are unchanged. They use the node's shared L1 provider, which derivation also uses, and a deadline there needs a separate decision.

It also proposes restructuring how the sequencer's blocks are signed:

  • The L1 watcher is the source of the unsafe block signer. Each L1 watcher chain owns a watch channel for the signer, seeded with the value read at startup. Gossip validation and block signing subscribe to it. In kona-node: recover from RPC failures without cascading across chains #23125 the L1 watcher sends each signer to the network actor over a bounded channel, and the network actor relays it into the gossip crate's own watch channel. Publishing to a watch channel cannot block, so the refresh loop no longer reserves channel capacity before sending.
  • Signing moves into a SignerActor between the sequencer and the network actor. kona-node: recover from RPC failures without cascading across chains #23125 signs in a task inside the network actor that gives each payload three 2s attempts and then skips it, so a transient signer outage leaves a gap in gossip. The SignerActor signs payloads in order and retries transient failures (remote signer transport errors and per-attempt timeouts) with backoff until they succeed, so nothing is skipped. Other signing errors are fatal, as before. If the signer rotates between signing and publishing, the payload is still published and peers reject it. The network actor goes back to only publishing.

Effect on operators:

  • A remote signer outage longer than the sequencer's 256-payload queue pauses block production until the signer recovers, and the sequencer keeps answering admin RPC throughout, so op-conductor can still stop it. The backlog is then gossiped in order. With kona-node: recover from RPC failures without cascading across chains #23125 the sequencer keeps producing blocks and skips gossip for blocks it failed to sign. A long outage is meant to page someone.
  • The block signer, including a remote signer's startup health check, starts only in sequencer mode. A validator with a signer configured logs that it is ignored.
  • A kona-node in sequencer mode now fails at startup without a block signer (--p2p.sequencer.key, --p2p.sequencer.key.path, or --p2p.signer.endpoint with --p2p.signer.address). Before, it started and dropped every payload, so its blocks were never gossiped. op-node still allows this, but there it is only useful with p2p disabled, which kona does not support. The kona-node sequencer docs and CLI reference now list the signer flags.

BREAKING CHANGE: a kona-node started with --mode sequencer and no block signer now fails to start; configure one of the signer flags above. The rest affects code built on the kona-gossip and kona-node-service crates.

  • GossipDriverBuilder::build returns only the driver, GossipDriverBuilder::with_unsafe_block_signer_receiver takes a watch::Receiver<Address>, and GossipDriverBuilder::new and GossipDriver::builder are no longer const.
  • NetworkBuilder::new loses its signer parameter, and NetworkBuilder::with_signer is removed. NetworkDriver and NetworkHandler lose signer and unsafe_block_signer_sender.
  • NetworkActor::new loses its signer-update receiver and takes a receiver of SignedPayload instead of unsigned payloads. L1WatcherChain::new takes a watch::Sender<Address>.
  • NetworkDriverError::BlockSignerStartError and NetworkActorError::{FailedToSignPayload, SigningTask, MissingUnsafeBlockSigner} are removed.
  • SignerActor::new takes a BlockSignerHandler; the actor always has a signer.
  • UnsafePayloadGossipClient has a new required method, has_capacity, which the sequencer checks before sealing.

To migrate: give the L1 watcher chain the sender of a signer watch channel and pass its receiver to NetworkBuilder::with_unsafe_block_signer. To gossip a sequencer's blocks, start the BlockSigner and run a SignerActor between the sequencer's payload channel and the network actor's publish channel. Callers with a fixed signer keep passing an Address to NetworkBuilder::new.

Reviews: rust-code-reviewer was run on each commit. For the transport commit, its main finding, that a hyper-level timeout misses a stalled response body, is fixed. deletion-reviewer was run on the signer commits. It found no surviving references to the removed APIs outside code, and its findings are applied: the breaking changes are marked, a dead error variant is removed, and the kona-node design doc lists the signer actor.

🤖 Generated with Claude Code

@joshklop
joshklop force-pushed the josh/kona-l1-signer-reloader branch from dcae0fc to d6de3c3 Compare October 5, 2026 20:12
@joshklop joshklop changed the title kona-node: reload the unsafe block signer on a timer kona-node: refresh the unsafe block signer only on new L1 heads Oct 5, 2026
@joshklop
joshklop force-pushed the josh/kona-l1-signer-reloader branch from 13bec91 to 9a3e40a Compare October 5, 2026 22:40
Replace the per-call `rpc_timeout` wrappers, the timeout inside
`record_call_time`, and the outer timeout around the reset forkchoice update
with one deadline on each of the engine client's RPC clients: a tower
`MapFutureLayer` on the alloy transport that times out each request. Every
request the engine client sends is now bounded, including ones that were not
wrapped, and a new call site cannot miss the deadline.

The deadline wraps alloy's transport future rather than the hyper client below
it: hyper's future completes when the response headers arrive, and alloy reads
the body afterwards. The new test covers both a request that is never answered
and a response whose body stalls after its headers.

The deadline is still per request, so a long reset traversal is not cut off,
and a timeout is still a transport error, so the engine retries it as before.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@joshklop joshklop changed the title kona-node: refresh the unsafe block signer only on new L1 heads kona-engine: bound engine client requests at the transport Oct 5, 2026
@joshklop
joshklop force-pushed the josh/kona-l1-signer-reloader branch from 9a3e40a to de042a8 Compare October 5, 2026 22:53
joshklop and others added 3 commits October 5, 2026 16:19
The L1 watcher sent each signer it read to the network actor over a bounded
channel, and the network actor relayed it into a watch channel that the gossip
crate created for block validation. Each L1 watcher chain now owns that watch
channel, seeded with the signer read at startup, and gossip validation
subscribes to it directly. The network actor's relay arm and its
closed-channel error are removed. Publishing a signer can no longer block, so
the refresh loop no longer reserves channel capacity and re-checks the head
before sending, and subscribers are only notified when the signer changes.

The gossip builder accepts an external signer receiver. Callers without an L1
watcher, such as `kona-node net` and the gossip example, still pass a fixed
address.

BREAKING CHANGE: affects code built on the `kona-gossip` and
`kona-node-service` crates. `GossipDriverBuilder::build` returns only the
driver, `GossipDriverBuilder::with_unsafe_block_signer_receiver` takes a
`watch::Receiver<Address>` instead of an `Address`, and
`GossipDriverBuilder::new` and `GossipDriver::builder` are no longer `const`.
`NetworkHandler` and `NetworkDriver` lose `unsafe_block_signer_sender`,
`NetworkActor::new` loses its signer-update receiver, and
`L1WatcherChain::new` takes a `watch::Sender<Address>`. To follow a changing
signer, create a watch channel, give its sender to the L1 watcher chain, and
pass its receiver to `NetworkBuilder::with_unsafe_block_signer`. To use a fixed
signer, keep passing an `Address` to `NetworkBuilder::new`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signing moves out of the network actor into a new `SignerActor` between the
sequencer and the network actor. It signs payloads in order and retries
transient signer failures (remote signer transport errors and per-attempt
timeouts) with capped backoff until they succeed, so no payload is skipped or
reordered. A long outage fills the sequencer's 256-payload queue and halts
block production. Other signing errors are fatal, as before. The expected
signer is read from the L1 watcher's watch channel on each attempt. If the
signer rotates between signing and publishing, the network actor publishes
the payload anyway and peers reject it.

This replaces the signing task inside the network actor, which gave each
payload three 2s attempts and then skipped it, leaving a gap in gossip.

The block signer is started only in sequencer mode, before the network binds
its ports. A sequencer without a signer drops its payloads with a warning, as
before. A validator with a signer configured logs that it is ignored, and no
longer health-checks a remote signer at startup.

BREAKING CHANGE: affects code built on the `kona-node-service` crate.
`NetworkBuilder::new` loses its `signer` parameter and
`NetworkBuilder::with_signer` is removed. `NetworkDriver` and
`NetworkHandler` lose `signer`. `NetworkDriverError::BlockSignerStartError`
and `NetworkActorError::{FailedToSignPayload, SigningTask,
MissingUnsafeBlockSigner}` are removed. `NetworkActor::new` takes a receiver
of `SignedPayload` instead of unsigned payloads. To gossip a sequencer's
blocks, start the `BlockSigner` and run a `SignerActor` between the
sequencer's payload channel and the network actor's publish channel.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The `SignerActor` now always holds a signer. A sequencer without one has
nothing to sign its blocks with, and its 256-payload queue would never be
read, so a node in sequencer mode now fails at startup when no block signer
is configured. Before, it started and dropped every payload with a warning,
so its blocks were never gossiped.

The kona-node sequencer docs and the CLI reference now list the signer flags,
and every sequencer example sets one.

BREAKING CHANGE: a kona-node started with `--mode sequencer` and none of
`--p2p.sequencer.key`, `--p2p.sequencer.key.path`, or
`--p2p.signer.endpoint` with `--p2p.signer.address` (or their
`KONA_NODE_P2P_*` environment variables) now fails to start. Configure one of
them. op-node still allows a sequencer without a signer. For code built on
`kona-node-service`, `SignerActor::new` takes a `BlockSignerHandler` instead
of an `Option`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@joshklop joshklop changed the title kona-engine: bound engine client requests at the transport kona-node!: bound engine RPCs at the transport and add a SignerActor Oct 5, 2026
@joshklop
joshklop marked this pull request as ready for review October 5, 2026 23:39
@joshklop
joshklop requested review from a team as code owners October 5, 2026 23:39
During a long signer outage the gossip queue fills. The sequencer sealed the
next block, committed it to op-conductor, and then waited for queue space
inside its step, so it stopped answering admin RPC, including op-conductor's
`StopSequencer`. A remote signer is typically shared by every sequencer in a
cluster, so failing over does not help, but the sequencer still has to answer
the conductor.

The sequencer now checks for queue space before it seals, and while the queue
is full it builds nothing and keeps answering admin queries. It is the queue's
only sender, so once there is room the hand-off after sealing cannot wait. No
block is skipped: building resumes on the next tick after the queue drains,
and the backlog is gossiped in order. The pause and resume are each logged
once.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
joshklop and others added 3 commits October 5, 2026 16:49
Log every build tick at info, including ticks skipped because the gossip
queue is full, instead of tracking the pause in a field to log only its start
and end.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A tick skipped because the gossip queue is full means block production is
paused, so log it at warn rather than info.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Move retries of transient signing failures from the SignerActor into the
remote signer, so the actor only sees errors that retrying cannot fix.

The remote signer's HTTP client now has a request timeout, which also bounds
the startup health check and clients rebuilt for new certificates. Its signing
request is retried with exponential backoff until it succeeds when the signer
is unreachable, does not answer in time, or answers with an error. A request
or response that cannot be encoded or decoded is not retried. The retry is in
the signing call rather than in transport middleware because the same client
makes the startup health check, which should fail fast.

The signer address is now checked once per payload instead of once per
attempt. If the unsafe block signer rotates while a request is being retried,
the signature returned is for the retired address and the next payload fails
with InvalidAddress, as a rotation already did.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant