Store rung V151: collect each Availability Group's group_id and count groups by it (#4475) - #4495
Merged
Merged
Conversation
…-history index (#4469, #4477) The startup watermark read looks up each collector's newest collection_log row through a new index instead of scanning the newest chunks: idx_collection_log_watermark on collection_log (server_id, collector_name, collection_time DESC), plus a per-collector unnest()/LATERAL LIMIT 1 lookup replacing the GROUP BY. Also adds idx_job_history_server_run on job_history (server_id, run_datetime DESC, instance_id DESC) for the Viewer's Job History read. Refs #4469 Refs #4477
…rm, the top-rung pins (#4469) StorageVersion.SchemaVersion moves to 150. The Viewer schema probe gains a V150 sentinel (both idx_collection_log_watermark and idx_job_history_server_run present) and its newest-first arm above V149's, so a fully-migrated store maps to exactly 150 instead of falling through to V149 and showing a spurious upgrade banner. QueryStoreLivenessHotTouchLiveTests' 'I am the top rung' facts move to a new CollectionLogWatermarkAndJobHistoryIndexesRungTests class; V149's and V148's own facts are rewritten to assert what stays true forever (registered, ladder-dense, gated below the current top's arm) rather than 'is exactly the top'. Refs #4469, #4477
AgTopology.CountDistinctGroups gains an overload that takes an optional group_id per member. A member carrying a group_id unions exactly on a matching id (case-insensitive); two members with different ids never union even when their names match and their replicas overlap; a member with no id falls back to the existing name-plus-replica-overlap rule among the other id-less members; and a with-id member still joins a same-named without-id member whose replicas overlap (the same AG seen before and after the group_id column landed on that reporter). The old two-tuple overload is kept for existing id-less callers and delegates to the new one with no id, so nothing that used it needs to change. group_id is threaded through every read that already carries the AG replica/database rows: AgTopology's row and card types, the Storage layer's DarlingAgStatesReader (appended last in both SELECTs and row constructors), the Viewer's AG mapping, the MCP/web AgReader's row and view types, and Lite's DuckDB read. A store or reader running before the group_id column exists never happens in practice, since the service migrates the store before reading it and the Viewer refuses a store below its required schema version — no fallback was added for that case.
Bumps CurrentSchemaVersion 64 -> 65 and adds an idempotent, non-fatal ALTER TABLE ... ADD COLUMN IF NOT EXISTS group_id VARCHAR for both ag_replica_states and ag_database_replica_states, matching the v62 cntr_type migration's shape. Without this the shared collectors' positional row writer fails Lite's AG collection the moment it appends the new trailing group_id value, since the DuckDB table would be one column short. Fresh installs already carry the column: the AG tables are generated from the shared collector catalog's PayloadColumns by DuckDbSchemaGenerator, not hand-written in Schema.cs, so no fresh-install schema edit is needed and DuckDbSchemaTests' table count is unaffected. Also fixes two pre-existing AgCollectorDefinitionTests call sites that predated the shared Row records' GroupId parameter and were failing to build.
…51-ag-group-id # Conflicts: # Darling/PerformanceMonitor.Darling.Storage/PgMigrations.cs # Darling/PerformanceMonitor.Darling.Storage/StorageVersion.cs # Darling/PerformanceMonitor.Darling.Viewer/ViewerDataService.cs
…ation round trip (#4475) AgGroupIdRungTests takes over the 'I am the top rung' facts for V151 from CollectionLogWatermarkAndJobHistoryIndexesRungTests, which now asserts what stays true forever (registered, ladder-dense, gated below the current top's arm) instead of 'is exactly the top'. QueryStoreLivenessHotTouchLiveTests' every-rung-above-me loop now zeros every ordinal from its own up through the newest, so it stays correct as new rungs land above it without a per-release edit. AgGroupIdRungTests also proves the live migration round trip on a compressed hypertable: climb to V150, plant a row in each of ag_replica_states and ag_database_replica_states, compress their chunks the product way, then apply V151. Both columns exist afterward, the pre-V151 rows read group_id IS NULL, and a closing INSERT in the collector's own column order carries a GUID string end to end. PgSchemaGeneratorTests' fresh-vs-upgraded shape pin for both AG collector tables now accounts for V151's appended group_id column on top of the existing V34/V36/V37 history, so it keeps proving a fresh store's generated schema matches an upgraded store's migrated one exactly.
…ungTests Timescale probe (#4475)
erikdarlingdata
marked this pull request as ready for review
September 27, 2026 20:06
erikdarlingdata
added a commit
that referenced
this pull request
Sep 28, 2026
…00Z UTC (#4621) Moves the CHANGELOG entries carried in merged pull-request descriptions into [Unreleased]. The cut is PRs merged after 2026-09-26T17:37:33Z and at or before 2026-09-28T17:40:00Z; the next splice starts after it. - 76 PRs are spliced: Fixed 36, Changed 28 and Added 12, each counted once under its first section. That is 79 bullets: #4481's Fixed entry holds three, and #4548 adds a second group under Added. - 21 PRs have no user-visible entry (None, test-only, CI-only, or deferred to a parent). - The [#4509] link definition, which #4538 also carries, is defined once. - The IMPORTANT upgrade notes from #4489, #4495, #4501, #4506 and #4541 are held for the release cut and are not in this change. - Only CHANGELOG.md changes.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #4475.
Why
Since #4481, the Availability Groups count joins the cards of the same AG by name plus overlapping replica sets. That leaves one case it can't see.
sys.dm_hadr_availability_replica_statesreturns only local information. So each secondary reports only itself, their replica sets share nothing, and they count as 2.group_idis the same GUID on every replica, so it identifies the group exactly.What changes
AgReplicaStatesCollectorandAgDatabaseReplicaStatesCollectorselectgroup_id = CONVERT(nvarchar(36), ag.group_id), appended as their last column.ALTER TABLE collect.ag_replica_states ADD COLUMN IF NOT EXISTS group_id text, and the same oncollect.ag_database_replica_states;StorageVersion.SchemaVersion= 151, and the Viewer's store probe gains V151's sentinel and its arm.DuckDbInitializer.CurrentSchemaVersion= 65 adds the same column to both tables, idempotently. Fresh Lite stores get it from the collectors' column lists.AgTopology.CountDistinctGroups):group_idare one group, even when they share no replica name;group_ids are never one group, even with the same name and overlapping replicas;group_idis carried through the store's AG state reader, the Viewer's AG view, the web/MCPget_ag_health(distinct_ag_count) and Lite's AG read.Test plan
AgTopologyCardsTests, pure):DarlingAgReader.BuildinDarlingAgReaderTests, with the id-less companion (→ 2).dev, the id-less form of the first case gives 2.AgGroupIdLiveReaderTests): two servers each report only their own replica of the same AG (the samegroup_id).DarlingAgReader.GetAgHealthAsyncgivesdistinct_ag_count= 1.GetAvailabilityGroupsAsync→AgTopology.Countsgives 1 group, 2 views, 2 reporting servers.devat runtime:Expected: 1 / Actual: 2.AgGroupIdRungTests, the new top rung):group_idNULL, and a write in the collector's column order carries the GUID.Lite.Testsbuilds, and Windows CI runs it.Darling.TestsandLite.Testsbuild with 0 warnings.CHANGELOG
SECTION: Changed
ENTRY:
group_id, so two monitored secondaries of one AG count as one group even without its primary ([Store rung V151: collect each Availability Group's group_id and count groups by it (#4475) #4495]) - store rung V151 adds the column. Rows collected before it fall back to matching by name and replica.REF:
[Store rung V151: collect each Availability Group's group_id and count groups by it (#4475) #4495]: Store rung V151: collect each Availability Group's group_id and count groups by it (#4475) #4495
IMPORTANT: Store rung V151 adds a
group_idcolumn toag_replica_statesandag_database_replica_states(Lite: schema 65). Rows collected before the upgrade have it empty.