Skip to content

benches: add grpc-rust (tonic-protobuf servers, grpc client) to the comparison suite - #293

Merged
iainmcgin merged 10 commits into
mainfrom
iain/bench-grpc-rust
Sep 26, 2026
Merged

iainmcgin merged 10 commits into
mainfrom
iain/bench-grpc-rust

Conversation

@iainmcgin

@iainmcgin iainmcgin commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator

grpc-rust's grpc crate is a client channel only, so the server-side comparison target is its tonic-protobuf codec (tonic over Google's protobuf v4 on upb), and the grpc channel itself can only be compared as a client. benches/rpc-grpc-rust ports the echo, log-ingest, fortunes and BenchService servers from benches/rpc-tonic to tonic-protobuf handler for handler (the log-ingest and fortunes servers return byte-identical responses to the buffa and prost ones), and adds client_bench for the client comparison.

The new crate is excluded from the workspace and built on demand by the drivers, because its codegen needs protoc 35.1 exactly plus grpc-rust's C++ protoc plugin, cmake-built on first use; CI never builds it, and task bench:grpc-rust:lint is its only lint. The existing drivers change too:

  • ServerProcess moves into rpc-bench's lib, RPC_BENCH_BIN_DIR points the drivers at prebuilt server binaries, and they panic if no request succeeds during warmup.
  • bench_server now serves through the generated BenchServiceServer dispatcher like the echo and log servers (--router restores the old path), which changes the connectrpc-rs latency column.
  • A unary_load driver is added, and echo_load takes a payload size.

The README Performance section is rewritten from a 2026-09 bare-metal run against tonic and tonic-protobuf. It replaces the 2026-03 table, which had tonic at 170.8 µs against 87.6 µs on small unary; the v0.2.0 harness re-run with tonic 0.14.5 or 0.14.6 shows parity, so that figure came from the original run's dependency set or host. Small unary, echo, and 10-message client and server streams are now within 3% across the three stacks. The raw-mode "~15% more capacity" claim is gone, because raw mode now measures within run-to-run noise of normal mode.

…n suite

grpc-rust (github.com/grpc/grpc-rust) ships `tonic-protobuf`, a tonic
codec over Google's `protobuf` v4 Rust runtime on the upb kernel. This
adds a bench crate with the echo, log-ingest, fortunes and BenchService
servers ported to that codec handler-for-handler from benches/rpc-tonic,
and wires them into cross_impl_bench, echo_bench, log_bench,
fortune_bench and profile_server.sh as a `tonic-protobuf` arm next to
`tonic` (prost) and `connectrpc-rs` (buffa).

The crate lives outside the cargo workspace and is built on demand by
the bench drivers, because its codegen needs protoc 35.1 exactly plus the
C++ protoc-gen-rust-grpc plugin (cmake-built from source on first use,
or supplied prebuilt via GRPC_RUST_PROTOC_DIR), and tonic-protobuf is
unpublished, so every grpc-rust crate is a git dependency pinned to one
revision. CI never builds it.

Driver changes that come with the new arm: the per-driver ServerProcess
copies move into rpc-bench's lib; every driver takes prebuilt server
binaries from RPC_BENCH_BIN_DIR when it is set, so the suite can run on
a benchmarking host with no toolchains; cross_impl_bench builds its
server set through grpc_servers() / connect_servers() instead of a
positional IMPLS array; and the load drivers assert that at least one
request succeeded during warmup, so a server that rejects every call
fails loudly instead of reporting 0 req/s.

Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>
…`grpc`)

grpc-rust's `grpc` crate is a client-side channel with no server, so the
only way to compare against it is a client benchmark. `client_bench`
holds the server constant (the connectrpc-rs echo server by default,
`--server-bin=PATH` to override) and drives it closed-loop with each
client stack in turn over gRPC/h2: connectrpc-rs on hyper-util's pooled
client, connectrpc-rs on its own Http2Connection, tonic's Channel with
prost stubs, and grpc-rust's grpc::client::Channel with grpc-protobuf
stubs on upb messages. `--conns=1,8` sweeps single-connection and
eight-connection configurations at concurrency 1, 16 and 64, reporting
req/s and p50/p99 from per-worker latency samples.

The tonic arm's stubs come from tonic-prost-build at the same grpc-rust
revision as the servers, fed the cmake-built protoc so the crate still
needs no system protoc; the connectrpc-rs arms reuse rpc-bench's
checked-in generated stubs through a path dependency.

Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>
…onic-protobuf

Re-runs the cross-impl latency, echo and log-ingest benchmarks on a
bare-metal c7i.metal-24xl (turbo disabled) against tonic 0.14.6 with
prost and with grpc-rust's tonic-protobuf (upb) codec, and rewrites the
Performance section and charts from that run. The previous figures were
from the 2026-03 initial-release run against tonic + prost only.

The picture changes: small-unary latency and echo throughput are now
level across the three stacks, where the 2026-03 table had tonic at
170.8 us against 87.6 us on small unary. That figure does not reproduce
with current dependencies: the v0.2.0 harness re-run today, with tonic
pinned to either 0.14.5 or 0.14.6, also measures tonic level with
connectrpc-rs, so the cause lies in the dependency versions or host of
the original run and was not isolated further. The buffa decode
advantage on the 50-record log batch is 38% over prost and 15% over upb
at concurrency 1, and 14% over prost / level with upb under load; upb is
14% slower on the 1 MB gzip payload; and connectrpc-rs is 9-12% slower
than both tonic configurations on the 10-message client stream. The
fortunes and CPU-profile subsections keep their 2026-03 numbers and now
say so.

Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>
…cho_load payload size

cross_impl_bench's bench_server served BenchService through the dynamic
Router while echo_server and log_server use the codegen
FooServiceServer<T> dispatcher the README describes; it now uses the
dispatcher too (--router keeps the old path for comparison; the
difference measured about 0.8 us per small unary call at c=1).
unary_load is a closed-loop BenchService.Unary driver with the
small_request() payload for profiling a server under perf or strace
without criterion in the picture, and echo_load takes an optional
payload size so a run can straddle h2's 256-byte DATA-frame chain
threshold.

Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>
@iainmcgin
iainmcgin force-pushed the iain/bench-grpc-rust branch from 49b2932 to 96550ed Compare September 6, 2026 15:24
…add the client-stacks table

The latency rows now come from bench_server on the codegen dispatcher,
matching the echo and log servers; every echo and log-ingest cell
reproduced within 1% on a second pass in the same session. The new
Client stacks subsection reports client_bench: the connectrpc-rs echo
server driven by the two connectrpc-rs transports, tonic's client and
grpc-rust's grpc channel, over 1 and 8 connections.

Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>
…ard CLI args

The multi-connection echo numbers are produced by `--multi-conn=8`, which
`task bench:echo -- --multi-conn=8` silently dropped because the task did
not forward CLI_ARGS. The Client stacks table states the protocol (gRPC
over h2 for all four clients), splits connections and requests in flight
into separate columns, drops the row that repeated the single-connection
case, and its summary sentence now matches the cells.

Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>
iainmcgin added a commit to allada/connect-rust that referenced this pull request Sep 24, 2026
…onnectrpc#294)

buffa 0.9.2 shipped on 2026-09-03 and changes codegen output (generated
`decode_view` attaches the element-memory budget, view decoders merge
split map values, generated enums carry
`#[allow(non_camel_case_types)]`), and because the workspace has no
committed lockfile the *Check generated code* job resolved 0.9.2
immediately and has failed on every PR since — connectrpc#293 is the first to show
it. This regenerates the five checked-in directories with the 0.9.2
plugins and raises the floor to match the regen baseline; code generated
by 0.9.1 still compiles against 0.9.2, so the floor is for the hardening
fixes and the baseline rather than a compile requirement. A second `task
generate:all` after this commit is a no-op.

Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>
Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>

# Conflicts:
#	.gitignore
…299)

Re-runs the latency, echo, log-ingest and client-stack benchmarks on a
bare-metal c7i.metal-24xl (turbo disabled, client and servers on one host
over loopback, unpinned) and rewrites those tables, the charts and the
number-tied prose from that run. The fortunes and CPU-profile subsections
keep their 2026-03 numbers, and the section header now says that the
2026-09 date applies unless a subsection says otherwise.

The 10-message client stream is now level across the three stacks
(166.3 us for connectrpc-rs against 168.2 us for tonic and 162.2 us for
tonic-protobuf), where the previous run had connectrpc-rs 9-14% slower;
the server now decodes streamed request messages in the handler's task,
so the summary drops the cross-task explanation. The log-batch figures
are now stated from the table's baseline (tonic 6-13% fewer requests per
second under load, tonic-protobuf within 4%), and the latency figure as
time taken (39% longer with prost, 16% with upb).

A second echo and log-ingest pass in the same session reproduced every
echo cell within 1.1% and every log-ingest cell within 1.4%, except
tonic-protobuf at c=256, which measured 23% lower on the second pass;
the README says so. The client-stacks prose no longer calls grpc-rust's
channel non-hyper, since its transport is built on hyper too, and the
cross bench's Go toolchain requirement is stated.

Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>
The row depends on what the bench process ran before it, and a
back-to-back run showed the previous table's client is not faster than
the current one.

Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>
Signed-off-by: Iain McGinniss <309153+iainmcgin@users.noreply.github.com>
@iainmcgin
iainmcgin marked this pull request as ready for review September 26, 2026 02:10
@iainmcgin
iainmcgin added this pull request to the merge queue Sep 26, 2026
Merged via the queue into main with commit d0429ac Sep 26, 2026
14 checks passed
@iainmcgin
iainmcgin deleted the iain/bench-grpc-rust branch September 26, 2026 02:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants