diff --git a/CHANGELOG.md b/CHANGELOG.md index d20e253d6..b74ba755e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -45,7 +45,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 - **The database-state heal tests no longer assume a best-effort maintenance cycle always runs** ([#2266] item 2, root-caused from source) - `RebaselinedByHandDuringAnOutage_HealsOnceTheDatabaseRecovers` failed on a PR whose diff could not reach it and passed on a re-run of the same commit, reporting `ExpectedState = SUSPECT` alongside `StateDesc = ONLINE`. That pair is only reachable one way: `GetDatabaseStateDeviationsAsync` performs its seeding, the [#2189] heal, the [#2203] forget and the prune inside a block that opens the write connection with a **5-second** lock acquisition and, on `TimeoutException`, skips the entire maintenance block while still running the deviation read - which is deliberate and documented, because skipping is the only lossless option when archival holds the lock. That write lock is **static, shared by the whole process**, and xunit runs test classes in parallel, so another class can hold it long enough for a cycle to skip its maintenance. The test was therefore asserting that the heal lands in ONE cycle, which the design does not promise; it now sweeps until the expectation settles, bounded, which is the actual contract. It cannot mask a regression: a genuinely broken heal never settles, every cycle runs, and the caller's own assertion fails on the final result with its own message exactly as before. Ruled out first: the fixture mints a unique temp directory per instance so no two classes share a store file, and the server id is used by no other class - the shared resource is the LOCK, not the data. **Not fixed here, and worth knowing**: that `catch (TimeoutException)` logs nothing, so in production a sustained contention window means baselines quietly stop being seeded and healed with no evidence anywhere, which matters because [#2189] exists precisely because an unhealed baseline inverts the alert permanently. Item 1 of that issue (the sub-second scale-test comparison) is deliberately untouched: every candidate fix needs the jitter distribution, and guessing a threshold is how an intermittent test stops looking broken without becoming correct. -- **`get_top_queries_by_cpu` can rank a procedure's dynamic SQL as ONE statement: `group_by: "host_object"`** ([#2235]) - `query_hash` is a SHAPE hash, so dynamic SQL built with per-value literals fragments one logical statement across as many hashes as there are literal sets. Measured on `prod-pos-use2-apex-01`: **21 hashes** for a single `API.GetInventoryWithLabsV5` `insert #result` statement, whose fragments were **58-65% of the instance's worker_time** in every window sampled - while the hash never entered the 168-hour top 20, and the per-query ranking as a whole accounted for roughly a **tenth** of the box's CPU. A top-N-by-hash list cannot surface that no matter how large N is, and nothing in the output said so. Setting `group_by: "host_object"` collapses every statement of a hosting procedure or function into one row, which is what makes the real consumer rank first; `distinct_query_hashes` reports how many hashes the row rolled up (21, in the reported case) and is the number that explains why the default ranking missed it, with a `rollup_note` saying so in words. `query_hash` and `query_text` in a rolled-up row are one representative fragment, exactly as `query_text` already is when `distinct_texts > 1`. **Ad-hoc statements keep their per-hash grouping in BOTH modes, and that is the load-bearing part**: ad-hoc rows carry `host_object_name = NULL`, so a bare `GROUP BY host_object_name` would pool every unrelated ad-hoc statement in a database into one meaningless row - a worse attribution bug than the one being fixed - and the representative-text lookup would start serving an unrelated statement's text. The grouping key keys those rows on their own `query_hash` instead, identical to the default read. The per-hash grouping ([#2012] stage 2) stays the DEFAULT and is unchanged: two procedures sharing a hash genuinely are different work, which is why that split exists, so this is an additional lens rather than a replacement. Implemented as a sibling SQL const rather than a built clause because Postgres cannot parameterize `GROUP BY` and every read here is a public const so the suite can pin its dialect without a live store; an unrecognised `group_by` is rejected rather than silently falling back, since a caller who asked for a rollup and got a per-hash ranking would read it as "this procedure is not hot", which is the exact wrong conclusion. `group_by` is trailing and optional, so Lite-shaped calls are unaffected, and it is deliberately absent from `get_top_procedures_by_cpu` (already keyed on the object) and `get_query_store_top` (keys on `query_id`, which does not fragment). Pinned by a live-Postgres test that asserts the collapse, the ad-hoc non-pooling, per-row text correctness, and that total CPU and executions are CONSERVED across both groupings - a rollup must redistribute attribution, never invent or lose it. +- **`get_top_queries_by_cpu` can rank a procedure's dynamic SQL as ONE statement: `group_by: "host_object"`** ([#2235]) - `query_hash` is a SHAPE hash, so dynamic SQL built with per-value literals fragments one logical statement across as many hashes as there are literal sets. Measured on `prod-sql-use2-alpha-01`: **21 hashes** for a single `API.GetInventoryWithLabsV5` `insert #result` statement, whose fragments were **58-65% of the instance's worker_time** in every window sampled - while the hash never entered the 168-hour top 20, and the per-query ranking as a whole accounted for roughly a **tenth** of the box's CPU. A top-N-by-hash list cannot surface that no matter how large N is, and nothing in the output said so. Setting `group_by: "host_object"` collapses every statement of a hosting procedure or function into one row, which is what makes the real consumer rank first; `distinct_query_hashes` reports how many hashes the row rolled up (21, in the reported case) and is the number that explains why the default ranking missed it, with a `rollup_note` saying so in words. `query_hash` and `query_text` in a rolled-up row are one representative fragment, exactly as `query_text` already is when `distinct_texts > 1`. **Ad-hoc statements keep their per-hash grouping in BOTH modes, and that is the load-bearing part**: ad-hoc rows carry `host_object_name = NULL`, so a bare `GROUP BY host_object_name` would pool every unrelated ad-hoc statement in a database into one meaningless row - a worse attribution bug than the one being fixed - and the representative-text lookup would start serving an unrelated statement's text. The grouping key keys those rows on their own `query_hash` instead, identical to the default read. The per-hash grouping ([#2012] stage 2) stays the DEFAULT and is unchanged: two procedures sharing a hash genuinely are different work, which is why that split exists, so this is an additional lens rather than a replacement. Implemented as a sibling SQL const rather than a built clause because Postgres cannot parameterize `GROUP BY` and every read here is a public const so the suite can pin its dialect without a live store; an unrecognised `group_by` is rejected rather than silently falling back, since a caller who asked for a rollup and got a per-hash ranking would read it as "this procedure is not hot", which is the exact wrong conclusion. `group_by` is trailing and optional, so Lite-shaped calls are unaffected, and it is deliberately absent from `get_top_procedures_by_cpu` (already keyed on the object) and `get_query_store_top` (keys on `query_id`, which does not fragment). Pinned by a live-Postgres test that asserts the collapse, the ad-hoc non-pooling, per-row text correctness, and that total CPU and executions are CONSERVED across both groupings - a rollup must redistribute attribution, never invent or lose it. - **The Query Store tick and the Query Store backfill no longer run against one server at the same time** ([#2165]) - the per-tick `query_store` collection and the [#2058] first-contact backfill were independent loops with no per-server coordination at all, and both do heavy Query Store text extraction. Dogfood evidence from a 4-core multi-tenant box mid-consolidation: a 64 MB backfill slice for a freshly restored database ran concurrently with the tick's collection of a SIBLING database - a 12:50:58 backfill ship overlapping a 12:51:09 tick completion - so roughly 128 MB of extraction was in flight at once on the box least able to afford it. That overlap is not bad luck: a big catalog arriving is exactly what triggers BOTH the backfill and budget-bound tick passes, so the two loops collide precisely when the server is already drowning. A per-server gate now excludes them, in both apps - Darling's `DarlingWorker` keyed by server id, Lite's `RemoteCollectorService` keyed by server - sharing one `QueryStoreServerGate` primitive beside `AbandonableStep` (that one bounds how long a step may hold a loop; this one bounds what may run beside it). **Nothing ever waits, deliberately.** Both sides try-acquire with a zero timeout and SKIP on failure, because these are shared fleet loops: an in-flight slice runs to a 180-300 second abandonment deadline, so a blocking acquire would let one slow server stall collection for the entire fleet - the [#2148] wedge arriving through a lock instead of a hang. **Skipping is safe for this collector specifically** because its window is a watermark ([#1960]): the next pass resumes from the same boundary, so a skipped pass defers rows rather than dropping them - which is also why the gate must not be reused for a collector whose window is wall-clock derived. The "tick wins, backfill defers" bias is realized by CADENCE rather than preemption (stopping a statement already running on the monitored server would mean killing it): the tick retries on its own ~1-minute interval against the backfill's 5, so it recovers five times faster from a collision, and a slice is byte-budgeted so it is short in the healthy case. The backfill takes the gate OUTSIDE its `AbandonableStep`, so an abandoned-but-still-wedged slice keeps the gate closed - the statement is genuinely still running on the server and the tick must keep yielding to it. Built on an interlocked flag rather than a `SemaphoreSlim`: a gate that never waits needs none of what a semaphore provides, and would otherwise own one undisposed kernel object per monitored server forever. Leases are idempotent on dispose, because a stray second `Dispose()` would otherwise clear a flag the OTHER loop had since taken and let both run at once - the exact condition being prevented, reached from the wrong direction. diff --git a/Darling/Darling.Tests/DarlingPeerDisclosureTests.cs b/Darling/Darling.Tests/DarlingPeerDisclosureTests.cs index 432ef5722..dba6fbefb 100644 --- a/Darling/Darling.Tests/DarlingPeerDisclosureTests.cs +++ b/Darling/Darling.Tests/DarlingPeerDisclosureTests.cs @@ -55,13 +55,13 @@ private static DarlingPeerDirectory.Snapshot TwoPeers() => { new PeerStoreConfig { - Name = "prod-pos-use2-monitor-01", + Name = "prod-sql-use2-monitor-01", Covers = "the readable replicas of those same 42 primaries, from us-east-2", Matches = { "use2" }, }, new PeerStoreConfig { - Name = "prod-pos-pg-monitor-01", + Name = "prod-sql-pg-monitor-01", Covers = "the Aurora PostgreSQL clusters", /* No matches: a peer that declares none is still disclosed, just never singled out. */ }, @@ -304,7 +304,7 @@ public void FromConfig_TrimsAndDropsEmptyEntriesAndBlankPatterns() fleet — the one normalization step that is a correctness fix rather than tidiness. */ Assert.Equal(new[] { "use2" }, peer.Matches); Assert.False(peer.CoversServerName("anything-at-all")); - Assert.True(peer.CoversServerName("prod-pos-USE2-apex-01")); + Assert.True(peer.CoversServerName("prod-sql-USE2-alpha-01")); } [Fact] @@ -314,18 +314,18 @@ public void CoversServerName_IsCaseInsensitiveSubstring_AndNeverTrueWithoutPatte var use2 = snapshot.Peers[0]; var postgres = snapshot.Peers[1]; - Assert.True(use2.CoversServerName("prod-pos-use2-ayr-01")); - Assert.True(use2.CoversServerName("PROD-POS-USE2-AYR-01")); - Assert.False(use2.CoversServerName("prod-pos-use1-ayr-01")); + Assert.True(use2.CoversServerName("prod-sql-use2-beta-01")); + Assert.True(use2.CoversServerName("PROD-SQL-USE2-BETA-01")); + Assert.False(use2.CoversServerName("prod-sql-use1-beta-01")); Assert.False(use2.CoversServerName(null)); Assert.False(use2.CoversServerName(" ")); /* No declared patterns means "cannot tell", which must never render as "yes". */ Assert.False(postgres.CoversServerName("anything")); - Assert.Equal(new[] { "prod-pos-use2-monitor-01" }, - snapshot.PeersCovering("prod-pos-use2-ayr-01").Select(p => p.Name)); - Assert.Empty(snapshot.PeersCovering("prod-pos-use1-ayr-01")); + Assert.Equal(new[] { "prod-sql-use2-monitor-01" }, + snapshot.PeersCovering("prod-sql-use2-beta-01").Select(p => p.Name)); + Assert.Empty(snapshot.PeersCovering("prod-sql-use1-beta-01")); } /* ───────────────────────── the MCP instructions ───────────────────────── */ @@ -343,8 +343,8 @@ public void Instructions_DiscloseThisStoreAndItsPeers_AboveTheToolCensus() var text = DarlingMcpInstructions.Build(TwoPeers()); Assert.Contains(Use1Covers, text, StringComparison.Ordinal); - Assert.Contains("prod-pos-use2-monitor-01 — the readable replicas", text, StringComparison.Ordinal); - Assert.Contains("prod-pos-pg-monitor-01 — the Aurora PostgreSQL clusters", text, StringComparison.Ordinal); + Assert.Contains("prod-sql-use2-monitor-01 — the readable replicas", text, StringComparison.Ordinal); + Assert.Contains("prod-sql-pg-monitor-01 — the Aurora PostgreSQL clusters", text, StringComparison.Ordinal); /* The point of the section is that a peer is NAMED, never contacted — say so where the agent will read it, or it will try to route a query at the sibling. */ @@ -382,7 +382,7 @@ private static JsonElement RenderedServerList(DarlingPeerDirectory.Snapshot peer { var rows = new List { - new(1, "prod-pos-use1-ayr-01", "ayr", 16, new DateTime(2026, 8, 19, 12, 0, 0, DateTimeKind.Utc)), + new(1, "prod-sql-use1-beta-01", "ayr", 16, new DateTime(2026, 8, 19, 12, 0, 0, DateTimeKind.Utc)), }; return JsonDocument @@ -400,7 +400,7 @@ public void ListServers_CarriesThePeerFleetsSummary() var fleets = root.GetProperty("peer_fleets").EnumerateArray().ToList(); Assert.Equal(2, fleets.Count); - Assert.Equal("prod-pos-use2-monitor-01", fleets[0].GetProperty("name").GetString()); + Assert.Equal("prod-sql-use2-monitor-01", fleets[0].GetProperty("name").GetString()); Assert.Equal("the readable replicas of those same 42 primaries, from us-east-2", fleets[0].GetProperty("covers").GetString()); Assert.Equal(new[] { "use2" }, fleets[0].GetProperty("matches").EnumerateArray().Select(m => m.GetString()).ToArray()); Assert.Empty(fleets[1].GetProperty("matches").EnumerateArray()); @@ -409,7 +409,7 @@ public void ListServers_CarriesThePeerFleetsSummary() /* The existing payload is untouched — the disclosure is additive here too. */ var server = Assert.Single(root.GetProperty("servers").EnumerateArray()); - Assert.Equal("prod-pos-use1-ayr-01", server.GetProperty("server_name").GetString()); + Assert.Equal("prod-sql-use1-beta-01", server.GetProperty("server_name").GetString()); Assert.Equal("ayr", server.GetProperty("display_name").GetString()); } @@ -439,8 +439,8 @@ no mention of the siblings is the strongest version of the wrong conclusion. */ var disclosure = DarlingPeerDirectory.EmptyRegistryDisclosure(TwoPeers()); Assert.Contains("one of SEVERAL monitoring this fleet", disclosure, StringComparison.Ordinal); - Assert.Contains("prod-pos-use2-monitor-01", disclosure, StringComparison.Ordinal); - Assert.Contains("prod-pos-pg-monitor-01", disclosure, StringComparison.Ordinal); + Assert.Contains("prod-sql-use2-monitor-01", disclosure, StringComparison.Ordinal); + Assert.Contains("prod-sql-pg-monitor-01", disclosure, StringComparison.Ordinal); Assert.Contains($"This store covers: {Use1Covers}.", disclosure, StringComparison.Ordinal); /* Unchanged with nothing declared, like every other surface. */ @@ -450,14 +450,14 @@ no mention of the siblings is the strongest version of the wrong conclusion. */ /* ───────────────────────── the resolution miss ───────────────────────── */ private const string MissWithoutPeers = - "Could not resolve server. Available servers:\nprod-pos-use1-ayr-01"; + "Could not resolve server. Available servers:\nprod-sql-use1-beta-01"; [Fact] public void ResolutionMiss_IsByteForByteUnchangedWithNothingDeclared() { var (resolved, error) = DarlingServerResolver.ResolveOrError( - new[] { Registered("prod-pos-use1-ayr-01") }, - "prod-pos-use2-ayr-01", + new[] { Registered("prod-sql-use1-beta-01") }, + "prod-sql-use2-beta-01", DarlingPeerDirectory.Snapshot.Empty); Assert.Equal(default, resolved); @@ -468,8 +468,8 @@ public void ResolutionMiss_IsByteForByteUnchangedWithNothingDeclared() public void ResolutionMiss_NamesThePeerWhoseDeclaredCoverageMatches() { var (resolved, error) = DarlingServerResolver.ResolveOrError( - new[] { Registered("prod-pos-use1-ayr-01") }, - "prod-pos-use2-ayr-01", + new[] { Registered("prod-sql-use1-beta-01") }, + "prod-sql-use2-beta-01", TwoPeers()); Assert.Equal(default, resolved); @@ -479,21 +479,21 @@ public void ResolutionMiss_NamesThePeerWhoseDeclaredCoverageMatches() the local list is still the right answer to the commonest miss (a typo). */ Assert.StartsWith(MissWithoutPeers, error, StringComparison.Ordinal); - Assert.Contains("'prod-pos-use2-ayr-01' is not monitored HERE", error, StringComparison.Ordinal); - Assert.Contains("matches the declared coverage of peer store prod-pos-use2-monitor-01", error, StringComparison.Ordinal); + Assert.Contains("'prod-sql-use2-beta-01' is not monitored HERE", error, StringComparison.Ordinal); + Assert.Contains("matches the declared coverage of peer store prod-sql-use2-monitor-01", error, StringComparison.Ordinal); Assert.Contains("That is a SEPARATE Darling store", error, StringComparison.Ordinal); Assert.Contains("this server cannot read it", error, StringComparison.Ordinal); Assert.Contains($"This store covers: {Use1Covers}.", error, StringComparison.Ordinal); /* The peer that declared no patterns must not be blamed for a name it never claimed. */ - Assert.DoesNotContain("prod-pos-pg-monitor-01", error, StringComparison.Ordinal); + Assert.DoesNotContain("prod-sql-pg-monitor-01", error, StringComparison.Ordinal); } [Fact] public void ResolutionMiss_WithNoMatchingPeer_ListsThemWithoutClaimingUnmonitored() { var (_, error) = DarlingServerResolver.ResolveOrError( - new[] { Registered("prod-pos-use1-ayr-01") }, + new[] { Registered("prod-sql-use1-beta-01") }, "some-other-box", TwoPeers()); @@ -503,8 +503,8 @@ public void ResolutionMiss_WithNoMatchingPeer_ListsThemWithoutClaimingUnmonitore /* Both peers are still disclosed: the declarations are prose plus optional patterns, not a live registry, so "no pattern matched" is not evidence the server is unmonitored. */ - Assert.Contains("prod-pos-use2-monitor-01", error, StringComparison.Ordinal); - Assert.Contains("prod-pos-pg-monitor-01", error, StringComparison.Ordinal); + Assert.Contains("prod-sql-use2-monitor-01", error, StringComparison.Ordinal); + Assert.Contains("prod-sql-pg-monitor-01", error, StringComparison.Ordinal); } [Fact] @@ -517,12 +517,12 @@ sentence must not say "That is a SEPARATE store" about a list of two. */ Stores = { new PeerStoreConfig { Name = "box2", Covers = "the replicas", Matches = { "use2" } }, - new PeerStoreConfig { Name = "box3", Covers = "the archive replicas", Matches = { "prod-pos" } }, + new PeerStoreConfig { Name = "box3", Covers = "the archive replicas", Matches = { "prod-sql" } }, }, }); var (_, error) = DarlingServerResolver.ResolveOrError( - new[] { Registered("prod-pos-use1-ayr-01") }, "prod-pos-use2-ayr-01", overlapping); + new[] { Registered("prod-sql-use1-beta-01") }, "prod-sql-use2-beta-01", overlapping); Assert.NotNull(error); Assert.Contains("these peer stores: box2 — the replicas; box3 — the archive replicas", error, StringComparison.Ordinal); @@ -536,7 +536,7 @@ public void ResolutionMiss_WithNoNameGiven_DisclosesPeersWithoutAccusingOne() /* The blank-name miss (several servers, no server_name passed) has no name to match, so the disclosure must list rather than accuse. */ var (_, error) = DarlingServerResolver.ResolveOrError( - new[] { Registered("prod-pos-use1-ayr-01"), Registered("prod-pos-use1-apex-01") }, + new[] { Registered("prod-sql-use1-beta-01"), Registered("prod-sql-use1-alpha-01") }, " ", TwoPeers()); @@ -564,8 +564,8 @@ public void ResolutionMiss_ReadsTheAmbientDeclaration_ThroughTheTwoArgOverload() Assert.Same(published.Snapshot, DarlingPeerDirectory.Current); var (_, error) = DarlingServerResolver.ResolveOrError( - new[] { Registered("prod-pos-use1-ayr-01") }, - "prod-pos-use2-ayr-01"); + new[] { Registered("prod-sql-use1-beta-01") }, + "prod-sql-use2-beta-01"); Assert.Contains("peer store box2 — the replicas", error, StringComparison.Ordinal); } @@ -577,8 +577,8 @@ public void ResolutionMiss_ReadsTheAmbientDeclaration_ThroughTheTwoArgOverload() Assert.True(DarlingPeerDirectory.Current.IsEmpty); var (_, afterReset) = DarlingServerResolver.ResolveOrError( - new[] { Registered("prod-pos-use1-ayr-01") }, - "prod-pos-use2-ayr-01"); + new[] { Registered("prod-sql-use1-beta-01") }, + "prod-sql-use2-beta-01"); Assert.Equal(MissWithoutPeers, afterReset); } diff --git a/Darling/Darling.Tests/HostObjectRollupLiveTests.cs b/Darling/Darling.Tests/HostObjectRollupLiveTests.cs index 713d7c650..a31567056 100644 --- a/Darling/Darling.Tests/HostObjectRollupLiveTests.cs +++ b/Darling/Darling.Tests/HostObjectRollupLiveTests.cs @@ -25,7 +25,7 @@ namespace PerformanceMonitor.Darling.Tests; /// /// The defect. query_hash is a SHAPE hash, so dynamic SQL built with per-value literals /// fragments one logical statement across as many hashes as there are literal sets — measured at 21 for a single -/// API.GetInventoryWithLabsV5 statement on prod-pos-use2-apex-01. Ranking by hash therefore +/// API.GetInventoryWithLabsV5 statement on prod-sql-use2-alpha-01. Ranking by hash therefore /// STRUCTURALLY cannot surface it: two fragments together were 58-65% of the instance's worker_time in every /// window sampled, while the hash never entered the 168-hour top 20 and the ranking as a whole accounted for /// roughly a tenth of the box's CPU. Nothing in the output said so. diff --git a/Darling/Darling.Tests/StoreIsOnThisMachineTests.cs b/Darling/Darling.Tests/StoreIsOnThisMachineTests.cs index ccde50e9d..505db6f8a 100644 --- a/Darling/Darling.Tests/StoreIsOnThisMachineTests.cs +++ b/Darling/Darling.Tests/StoreIsOnThisMachineTests.cs @@ -79,7 +79,7 @@ public void AnAbsentConnectionStringIsSilent(string? connectionString) /// THE CASE THAT WARNS: a store on another host, which is how the #2255 report was configured. [Theory] - [InlineData("Host=prod-pos-use2-monitor-01;Port=5641;Database=darling")] + [InlineData("Host=prod-sql-use2-monitor-01;Port=5641;Database=darling")] [InlineData("Host=10.149.55.242;Port=5432;Database=darling")] [InlineData("Host=store.internal.example.com;Database=darling")] [InlineData("Server=otherbox;Database=darling")] diff --git a/Darling/Darling.Tests/SweepPressureClassifierTests.cs b/Darling/Darling.Tests/SweepPressureClassifierTests.cs index bbc6b5fa6..39fb27e8c 100644 --- a/Darling/Darling.Tests/SweepPressureClassifierTests.cs +++ b/Darling/Darling.Tests/SweepPressureClassifierTests.cs @@ -18,7 +18,7 @@ namespace Darling.Tests; /// SKUs' get_collection_health serve so half-rate collection stops being visible only as a service-log /// warning. This SAME table is pinned identically in Lite.Tests so the two SKUs cannot drift. /// -/// The load-bearing case is the motivating measurement: prod-pos-use2-multi-01's four heavy +/// The load-bearing case is the motivating measurement: prod-sql-use2-multi-01's four heavy /// collectors averaged 22,141 + 16,590 + 13,544 + 8,437 ms against a 60s cadence — the body could not /// fit, every relaunch was skipped (~50 warnings/hour), the server collected at half rate, and all 40 /// collectors read HEALTHY, because from each one's own seat nothing was wrong. diff --git a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingPeerDirectory.cs b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingPeerDirectory.cs index 8c7fa48fa..c99080835 100644 --- a/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingPeerDirectory.cs +++ b/Darling/PerformanceMonitor.Darling.Service/Mcp/DarlingPeerDirectory.cs @@ -234,7 +234,7 @@ internal static string ResolutionMissDisclosure(Snapshot snapshot, string? reque if (matching.Count > 0) { /* Two peers can legitimately both claim a name (overlapping `matches`, e.g. "use1" and - "prod-pos"), so the follow-on sentence agrees in number rather than saying "That is a SEPARATE + "prod-sql"), so the follow-on sentence agrees in number rather than saying "That is a SEPARATE store" about a list of two. */ var single = matching.Count == 1; text.Append(subject) diff --git a/Darling/PerformanceMonitor.Darling.Service/darling.sample.json b/Darling/PerformanceMonitor.Darling.Service/darling.sample.json index 88b533d38..f8dbd339c 100644 --- a/Darling/PerformanceMonitor.Darling.Service/darling.sample.json +++ b/Darling/PerformanceMonitor.Darling.Service/darling.sample.json @@ -291,12 +291,12 @@ // "thisStoreCovers": "the 42 us-east-1 SQL Server primaries", // "stores": [ // { - // "name": "prod-pos-use2-monitor-01", + // "name": "prod-sql-use2-monitor-01", // "covers": "the readable replicas of those same 42 primaries, in-region from us-east-2", // "matches": ["use2"] // }, // { - // "name": "prod-pos-pg-monitor-01", + // "name": "prod-sql-pg-monitor-01", // "covers": "the Aurora PostgreSQL clusters", // "matches": ["-aurora-", ".cluster-"] // } diff --git a/Darling/README.md b/Darling/README.md index 97ad4e15e..826747eb4 100644 --- a/Darling/README.md +++ b/Darling/README.md @@ -510,11 +510,11 @@ Once enabled, open `http://localhost:5153/` in a browser on the service host. Li "thisStoreCovers": "the 42 us-east-1 SQL Server primaries", "stores": [ { - "name": "prod-pos-use2-monitor-01", + "name": "prod-sql-use2-monitor-01", "covers": "the readable replicas of those same 42 primaries, in-region from us-east-2", "matches": ["use2"] }, - { "name": "prod-pos-pg-monitor-01", "covers": "the Aurora PostgreSQL clusters", "matches": ["-aurora-"] } + { "name": "prod-sql-pg-monitor-01", "covers": "the Aurora PostgreSQL clusters", "matches": ["-aurora-"] } ] } ``` diff --git a/Lite.Tests/SweepPressureClassifierTests.cs b/Lite.Tests/SweepPressureClassifierTests.cs index 379b94f18..e34baec0f 100644 --- a/Lite.Tests/SweepPressureClassifierTests.cs +++ b/Lite.Tests/SweepPressureClassifierTests.cs @@ -18,7 +18,7 @@ namespace Lite.Tests; /// SKUs' get_collection_health serve so half-rate collection stops being visible only as a service-log /// warning. This SAME table is pinned identically in Darling.Tests so the two SKUs cannot drift. /// -/// The load-bearing case is the motivating measurement: prod-pos-use2-multi-01's four heavy +/// The load-bearing case is the motivating measurement: prod-sql-use2-multi-01's four heavy /// collectors averaged 22,141 + 16,590 + 13,544 + 8,437 ms against a 60s cadence — the body could not /// fit, every relaunch was skipped (~50 warnings/hour), the server collected at half rate, and all 40 /// collectors read HEALTHY, because from each one's own seat nothing was wrong.