Skip to content

Commit f4615e2

Browse files
committed
chore(docs): clarify the supported gRPC network boundary
Signed-off-by: Yordis Prieto <yordis.prieto@gmail.com>
1 parent aecb79b commit f4615e2

15 files changed

Lines changed: 186 additions & 156 deletions

‎.gitignore‎

Lines changed: 0 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -33,8 +33,6 @@ ipch/
3333
.DS_Store
3434

3535
src/EventStore/EventStore.Common/Properties/AssemblyVersion.cs
36-
src/EventStore/EventStore.ClientAPI/Properties/AssemblyVersion.cs
37-
3836
*.o
3937
*.ii
4038
*.s

‎Dockerfile‎

Lines changed: 1 addition & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -42,7 +42,6 @@ FROM mcr.microsoft.com/dotnet/sdk:10.0-${CONTAINER_RUNTIME} AS test
4242
WORKDIR /build
4343
COPY --from=build ./build/published-tests ./published-tests
4444
COPY --from=build ./build/ci ./ci
45-
COPY --from=build ./build/src/EventStore.Core.Tests/Services/Transport/Tcp/test_certificates/ca/ca.crt /usr/local/share/ca-certificates/ca_eventstore_test.crt
4645
COPY ./scripts/test.sh /build/test.sh
4746
RUN mkdir ./test-results
4847
RUN chmod +x /build/test.sh
@@ -89,7 +88,7 @@ ReplicationIp: 0.0.0.0" >> /etc/eventstore/eventstore.conf
8988

9089
VOLUME /var/lib/eventstore /var/log/eventstore
9190

92-
EXPOSE 1112/tcp 1113/tcp 2113/tcp
91+
EXPOSE 1112/tcp 2113/tcp
9392

9493
HEALTHCHECK --interval=5s --timeout=5s --retries=24 \
9594
CMD curl --fail --insecure https://localhost:2113/-/liveness || curl --fail http://localhost:2113/-/liveness || exit 1

‎docker-compose.yml‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -17,6 +17,7 @@ services:
1717
- shared.env
1818
environment:
1919
- EVENTSTORE_GOSSIP_SEED=172.30.240.12:2113,172.30.240.13:2113
20+
- EVENTSTORE_NODE_IP=172.30.240.11
2021
- EVENTSTORE_REPLICATION_IP=172.30.240.11
2122
- EVENTSTORE_CERTIFICATE_FILE=/etc/eventstore/certs/node/node.crt
2223
- EVENTSTORE_CERTIFICATE_PRIVATE_KEY_FILE=/etc/eventstore/certs/node/node.key
@@ -44,6 +45,7 @@ services:
4445
- shared.env
4546
environment:
4647
- EVENTSTORE_GOSSIP_SEED=172.30.240.11:2113,172.30.240.13:2113
48+
- EVENTSTORE_NODE_IP=172.30.240.12
4749
- EVENTSTORE_REPLICATION_IP=172.30.240.12
4850
- EVENTSTORE_CERTIFICATE_FILE=/etc/eventstore/certs/node/node.crt
4951
- EVENTSTORE_CERTIFICATE_PRIVATE_KEY_FILE=/etc/eventstore/certs/node/node.key
@@ -71,6 +73,7 @@ services:
7173
- shared.env
7274
environment:
7375
- EVENTSTORE_GOSSIP_SEED=172.30.240.11:2113,172.30.240.12:2113
76+
- EVENTSTORE_NODE_IP=172.30.240.13
7477
- EVENTSTORE_REPLICATION_IP=172.30.240.13
7578
- EVENTSTORE_CERTIFICATE_FILE=/etc/eventstore/certs/node/node.crt
7679
- EVENTSTORE_CERTIFICATE_PRIVATE_KEY_FILE=/etc/eventstore/certs/node/node.key

‎docs/README.md‎

Lines changed: 8 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -10,9 +10,13 @@ surfaces for running a node or cluster.
1010

1111
TrogonEventStore keeps the database node focused on the durable event log:
1212

13-
- Application event access is gRPC-first.
14-
- HTTP is reserved for the Admin UI, health probes, metrics, and other
15-
infrastructure-level concerns.
13+
- Database client APIs, cluster coordination, and follower-to-leader forwarding
14+
use gRPC over the node HTTP(S) endpoint. Database replication uses gRPC over
15+
a dedicated replication HTTP(S) endpoint.
16+
- Regular HTTP routes are reserved for the Admin UI, health probes, metrics,
17+
and other infrastructure-level concerns.
18+
- The server does not open a separate legacy EventStore TCP protocol listener
19+
or support its TCP transport configuration.
1620
- The project is FOSS-only. The documentation does not describe unsupported
1721
proprietary server features.
1822
- Rich read models, user-defined query engines, connector runtimes, and
@@ -37,7 +41,7 @@ For a production node, review:
3741

3842
## Protocols and clients
3943

40-
The supported application protocol is gRPC. Existing TrogonEventStore-compatible
44+
The supported database protocol is gRPC. Existing TrogonEventStore-compatible
4145
gRPC clients can be useful while the TrogonDB client libraries continue to
4246
evolve, but the server documentation should be treated as authoritative for this
4347
repository.

‎docs/admin-ui.md‎

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -10,7 +10,7 @@ The TrogonEventStore Admin UI is available at `http://SERVER_IP:2113/ui` and hel
1010

1111
The dashboard opens at `/ui` and combines the daily operational view in one place:
1212

13-
- _Cluster status_: live gossip membership, node state, checkpoints, TCP and HTTP endpoints, replica status, and a copy-friendly snapshot.
13+
- _Cluster status_: live gossip membership, node state, checkpoints, HTTP(S) endpoints, replica status, and a copy-friendly snapshot.
1414
- _Queue pressure_: live queue length, throughput, processing time, and currently processed messages.
1515
- _Node probes_: inline Ping, Node info, and Gossip checks rendered inside the UI.
1616

@@ -24,7 +24,8 @@ The _Observability_ page focuses on runtime diagnostics:
2424

2525
- queue groups and individual queue rows
2626
- current and last processed messages
27-
- TCP connection statistics
27+
- active connections on the node and replication HTTP/gRPC endpoints, including client identity, protocol, security, traffic rates, totals, and pending bytes
28+
- live gRPC replication sessions, byte totals, pending bytes, and send queue depth
2829
- snapshot output for copy-paste debugging
2930

3031
## Configuration

‎docs/architecture.md‎

Lines changed: 11 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -22,6 +22,17 @@ and sharding work. Features that need their own compute model, query model, or
2222
serving model should be separate components that consume the database through
2323
subscriptions or reads.
2424

25+
## Network protocol boundary
26+
27+
gRPC carries database client APIs, cluster coordination, and follower-to-leader
28+
forwarding over the node HTTP(S) endpoint. Ordinary HTTP routes on that endpoint
29+
serve the Admin UI, health probes, metrics, and other operator workflows.
30+
Database replication uses gRPC over a dedicated replication HTTP(S) endpoint so
31+
operators can isolate high-volume replication traffic from clients.
32+
33+
Both endpoints use the same TLS and node identity configuration. The node has no
34+
separate legacy EventStore TCP protocol listener or configuration surface.
35+
2536
## Projection execution
2637

2738
Projection execution is future external component work by default.

‎docs/cluster.md‎

Lines changed: 13 additions & 5 deletions
Original file line numberDiff line numberDiff line change
@@ -55,15 +55,22 @@ The multi-address DNS name cluster discovery only works for clusters that use ce
5555

5656
### Internal communication
5757

58-
When setting up a cluster, the nodes must be able to reach each other over both the HTTP channel, and the internal TCP channel. You should ensure that these ports are open on firewalls on the machines and between the machines.
58+
Cluster nodes use gRPC over each node's dedicated replication HTTP(S) endpoint
59+
for database replication. Gossip, elections, and follower-to-leader request
60+
forwarding use the node HTTP(S) endpoint. Ensure every node can reach both
61+
advertised endpoints on every other node.
5962

60-
Learn more about [internal TCP configuration](networking.md#replication-protocol) and [HTTP configuration](networking.md#http-configuration) to set up the cluster properly.
63+
Learn more about the [node](networking.md#http-configuration) and
64+
[replication](networking.md#internal-cluster-traffic) endpoints before
65+
configuring cluster firewall or network policy rules.
6166

6267
## Cluster with DNS
6368

6469
When you tell TrogonEventStore to use DNS for its gossip, the server will resolve the DNS name to a list of IP addresses and connect to each of those addresses to find other nodes. This method is very flexible because you can change the list of nodes on your DNS server without changing the cluster configuration. The DNS method is also useful in automated deployment scenarios when you control both the cluster deployment and the DNS server from your infrastructure-as-code scripts.
6570

66-
To use DNS discovery, you need to set the `ClusterDns` option to the DNS name that allows making an HTTP call to it. When the server starts, it will attempt to make a gRPC call using the `https://<cluster-dns>:<gossip-port>` URL (`http` if the cluster is insecure).
71+
To use DNS discovery, set the `ClusterDns` option to a DNS name that resolves to
72+
the cluster nodes. When the server starts, it attempts a gRPC call over
73+
`https://<cluster-dns>:<gossip-port>` (`http` if the cluster is insecure).
6774

6875
When using a certificate signed by a publicly trusted CA, you'd normally use the wildcard certificate. Ensure that the cluster DNS name fits the wildcard, otherwise the request will fail on SSL check.
6976

@@ -107,7 +114,8 @@ The setting accepts a comma-separated list of IP addresses or host names with th
107114

108115
TrogonEventStore uses a quorum-based replication model. When working normally, a cluster has one node known as a leader, and the remaining nodes are followers. The leader node is responsible for coordinating writes while it is the leader. Cluster nodes use a consensus algorithm to determine which node should be the leader and which should be followers. TrogonEventStore bases the decision as to which node should be the leader on a number of factors.
109116

110-
For a cluster node to have this information available to them, the nodes gossip with other nodes in the cluster. Gossip runs over HTTP interfaces of cluster nodes.
117+
For a cluster node to have this information available to them, the nodes gossip
118+
with other nodes in the cluster over the node HTTP(S) endpoint.
111119

112120
The gossip protocol configuration can be changed using the settings listed below. Pay attention to the settings related to time, like intervals and timeouts, when running in a cloud environment.
113121

@@ -229,7 +237,7 @@ candidate.
229237

230238
### Follower
231239

232-
A cluster assigns the follower role based on an election process. A cluster uses one or more nodes with the follower role to form the quorum, or the majority of nodes necessary to confirm that the write is persisted.
240+
A cluster assigns the follower role based on an election process. A cluster uses one or more nodes with the follower role to form the quorum, or the majority of nodes necessary to confirm that the write is persisted. When a follower accepts a request that must run on the leader, it forwards the request to the leader over gRPC on the leader's HTTP(S) endpoint.
233241

234242
### Read-only replica
235243

‎docs/diagnostics/README.md‎

Lines changed: 17 additions & 12 deletions
Original file line numberDiff line numberDiff line change
@@ -6,6 +6,7 @@ TrogonEventStore provides several ways to diagnose and troubleshoot issues.
66
- [Metrics](metrics.md): collect standard metrics using Prometheus or OpenTelemetry.
77
- [Monitoring and alerting](monitoring.md): turn health, metrics, and logs into operational signals.
88
- [Stats](#statistics): runtime statistics exposed through the monitoring gRPC service.
9+
- [Connection statistics](#connection-statistics): active HTTP and gRPC connections on both listeners.
910

1011
You can also use external tools to measure the performance of TrogonEventStore and monitor the cluster health. Learn more on the [Integrations](./integrations.md) page.
1112

@@ -19,6 +20,21 @@ cluster-wide state.
1920
`Monitoring.Stats` collects fresh node-local statistics and memoizes the result for up to one second. It does
2021
not read persisted statistics events.
2122

23+
## Connection statistics
24+
25+
Use the `Monitoring.ConnectionStats` RPC to inspect active connections on the
26+
node and replication HTTP(S) listeners. Each result identifies the local and
27+
remote endpoints, connection ID, observed client name, protocol, application,
28+
TLS state, connection time, total bytes, and pending bytes.
29+
30+
Use `Monitoring.ReplicationStats` when the diagnostic question is specifically
31+
about live database replication sessions. Its results include the subscription
32+
and connection IDs, peer endpoint, byte totals, pending bytes, and send queue
33+
depth.
34+
35+
Both RPCs report node-local snapshots. Query each cluster member when diagnosing
36+
a cluster-wide connection or replication problem.
37+
2238
When statistics persistence is enabled, each node writes events to a reserved `$stats-<host:port>` stream. For
2339
example, a single local node writes to `$stats-127.0.0.1:2113`.
2440

@@ -49,17 +65,6 @@ type `$statsCollected`.
4965
"proc-gc-largeHeapSize": 0,
5066
"proc-gc-timeInGc": 0.0,
5167
"proc-gc-totalBytesInHeaps": 0,
52-
"proc-tcp-connections": 0,
53-
"proc-tcp-receivingSpeed": 0.0,
54-
"proc-tcp-sendingSpeed": 0.0,
55-
"proc-tcp-inSend": 0,
56-
"proc-tcp-measureTime": "00:00:19.0534210",
57-
"proc-tcp-pendingReceived": 0,
58-
"proc-tcp-pendingSend": 0,
59-
"proc-tcp-receivedBytesSinceLastRun": 0,
60-
"proc-tcp-receivedBytesTotal": 0,
61-
"proc-tcp-sentBytesSinceLastRun": 0,
62-
"proc-tcp-sentBytesTotal": 0,
6368
"es-checksum": 1613144,
6469
"es-checksumNonFlushed": 1613144,
6570
"sys-drive-/System/Volumes/Data-availableBytes": 545628151808,
@@ -104,7 +109,7 @@ type `$statsCollected`.
104109
"es-queue-MonitoringQueue-lengthLifetimePeak": 0,
105110
"es-queue-MonitoringQueue-totalItemsProcessed": 14,
106111
"es-queue-MonitoringQueue-inProgressMessage": "<none>",
107-
"es-queue-MonitoringQueue-lastProcessedMessage": "GetFreshTcpConnectionStats",
112+
"es-queue-MonitoringQueue-lastProcessedMessage": "GetFreshStats",
108113
"es-queue-PersistentSubscriptions-queueName": "PersistentSubscriptions",
109114
"es-queue-PersistentSubscriptions-groupName": "",
110115
"es-queue-PersistentSubscriptions-avgItemsPerSecond": 1,

‎docs/installation.md‎

Lines changed: 3 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -76,6 +76,9 @@ Before running a durable node or cluster:
7676
- Store data, index, and logs on durable volumes.
7777
- Expose `/-/liveness`, `/-/readiness`, and `/-/metrics` to the platform.
7878
- Use gRPC clients for application reads and writes.
79+
- Expose the node HTTP(S) endpoint to clients and operators. Allow peer nodes to
80+
reach the dedicated replication HTTP(S) endpoint on the private cluster
81+
network. No legacy EventStore TCP protocol listener is required or supported.
7982

8083
## Linux service notes
8184

0 commit comments

Comments
 (0)