You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: docs/admin-ui.md
+3-2Lines changed: 3 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -10,7 +10,7 @@ The TrogonEventStore Admin UI is available at `http://SERVER_IP:2113/ui` and hel
10
10
11
11
The dashboard opens at `/ui` and combines the daily operational view in one place:
12
12
13
-
-_Cluster status_: live gossip membership, node state, checkpoints, TCP and HTTP endpoints, replica status, and a copy-friendly snapshot.
13
+
-_Cluster status_: live gossip membership, node state, checkpoints, HTTP(S) endpoints, replica status, and a copy-friendly snapshot.
14
14
-_Queue pressure_: live queue length, throughput, processing time, and currently processed messages.
15
15
-_Node probes_: inline Ping, Node info, and Gossip checks rendered inside the UI.
16
16
@@ -24,7 +24,8 @@ The _Observability_ page focuses on runtime diagnostics:
24
24
25
25
- queue groups and individual queue rows
26
26
- current and last processed messages
27
-
- TCP connection statistics
27
+
- active connections on the node and replication HTTP/gRPC endpoints, including client identity, protocol, security, traffic rates, totals, and pending bytes
28
+
- live gRPC replication sessions, byte totals, pending bytes, and send queue depth
Copy file name to clipboardExpand all lines: docs/cluster.md
+13-5Lines changed: 13 additions & 5 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -55,15 +55,22 @@ The multi-address DNS name cluster discovery only works for clusters that use ce
55
55
56
56
### Internal communication
57
57
58
-
When setting up a cluster, the nodes must be able to reach each other over both the HTTP channel, and the internal TCP channel. You should ensure that these ports are open on firewalls on the machines and between the machines.
58
+
Cluster nodes use gRPC over each node's dedicated replication HTTP(S) endpoint
59
+
for database replication. Gossip, elections, and follower-to-leader request
60
+
forwarding use the node HTTP(S) endpoint. Ensure every node can reach both
61
+
advertised endpoints on every other node.
59
62
60
-
Learn more about [internal TCP configuration](networking.md#replication-protocol) and [HTTP configuration](networking.md#http-configuration) to set up the cluster properly.
63
+
Learn more about the [node](networking.md#http-configuration) and
64
+
[replication](networking.md#internal-cluster-traffic) endpoints before
65
+
configuring cluster firewall or network policy rules.
61
66
62
67
## Cluster with DNS
63
68
64
69
When you tell TrogonEventStore to use DNS for its gossip, the server will resolve the DNS name to a list of IP addresses and connect to each of those addresses to find other nodes. This method is very flexible because you can change the list of nodes on your DNS server without changing the cluster configuration. The DNS method is also useful in automated deployment scenarios when you control both the cluster deployment and the DNS server from your infrastructure-as-code scripts.
65
70
66
-
To use DNS discovery, you need to set the `ClusterDns` option to the DNS name that allows making an HTTP call to it. When the server starts, it will attempt to make a gRPC call using the `https://<cluster-dns>:<gossip-port>` URL (`http` if the cluster is insecure).
71
+
To use DNS discovery, set the `ClusterDns` option to a DNS name that resolves to
72
+
the cluster nodes. When the server starts, it attempts a gRPC call over
73
+
`https://<cluster-dns>:<gossip-port>` (`http` if the cluster is insecure).
67
74
68
75
When using a certificate signed by a publicly trusted CA, you'd normally use the wildcard certificate. Ensure that the cluster DNS name fits the wildcard, otherwise the request will fail on SSL check.
69
76
@@ -107,7 +114,8 @@ The setting accepts a comma-separated list of IP addresses or host names with th
107
114
108
115
TrogonEventStore uses a quorum-based replication model. When working normally, a cluster has one node known as a leader, and the remaining nodes are followers. The leader node is responsible for coordinating writes while it is the leader. Cluster nodes use a consensus algorithm to determine which node should be the leader and which should be followers. TrogonEventStore bases the decision as to which node should be the leader on a number of factors.
109
116
110
-
For a cluster node to have this information available to them, the nodes gossip with other nodes in the cluster. Gossip runs over HTTP interfaces of cluster nodes.
117
+
For a cluster node to have this information available to them, the nodes gossip
118
+
with other nodes in the cluster over the node HTTP(S) endpoint.
111
119
112
120
The gossip protocol configuration can be changed using the settings listed below. Pay attention to the settings related to time, like intervals and timeouts, when running in a cloud environment.
113
121
@@ -229,7 +237,7 @@ candidate.
229
237
230
238
### Follower
231
239
232
-
A cluster assigns the follower role based on an election process. A cluster uses one or more nodes with the follower role to form the quorum, or the majority of nodes necessary to confirm that the write is persisted.
240
+
A cluster assigns the follower role based on an election process. A cluster uses one or more nodes with the follower role to form the quorum, or the majority of nodes necessary to confirm that the write is persisted. When a follower accepts a request that must run on the leader, it forwards the request to the leader over gRPC on the leader's HTTP(S) endpoint.
Copy file name to clipboardExpand all lines: docs/diagnostics/README.md
+17-12Lines changed: 17 additions & 12 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -6,6 +6,7 @@ TrogonEventStore provides several ways to diagnose and troubleshoot issues.
6
6
-[Metrics](metrics.md): collect standard metrics using Prometheus or OpenTelemetry.
7
7
-[Monitoring and alerting](monitoring.md): turn health, metrics, and logs into operational signals.
8
8
-[Stats](#statistics): runtime statistics exposed through the monitoring gRPC service.
9
+
-[Connection statistics](#connection-statistics): active HTTP and gRPC connections on both listeners.
9
10
10
11
You can also use external tools to measure the performance of TrogonEventStore and monitor the cluster health. Learn more on the [Integrations](./integrations.md) page.
11
12
@@ -19,6 +20,21 @@ cluster-wide state.
19
20
`Monitoring.Stats` collects fresh node-local statistics and memoizes the result for up to one second. It does
20
21
not read persisted statistics events.
21
22
23
+
## Connection statistics
24
+
25
+
Use the `Monitoring.ConnectionStats` RPC to inspect active connections on the
26
+
node and replication HTTP(S) listeners. Each result identifies the local and
0 commit comments