From a6b5ea69417ad1d147ab0fed789bb54e8573c574 Mon Sep 17 00:00:00 2001 From: Erik Darling <2136037+erikdarlingdata@users.noreply.github.com> Date: Wed, 30 Sep 2026 14:37:26 -0400 Subject: [PATCH 1/2] CHANGELOG: splice the 09-28 -> 09-30 merges into [Unreleased] Splices the entries carried by PRs merged to dev after 2026-09-28T17:40:00Z: 4 Added, 18 Changed and 120 Fixed bullets, with 104 definitions. Folds the daylight-saving and alert-retry families into one entry each, combines #4739 with #4763, replaces #4632's and #4688's row-label text with #4698's end state, and amends the [Unreleased] entries for #4278, #4341, #4396 and #4413. --- CHANGELOG.md | 254 ++++++++++++++++++++++++++++++++++++++++++++++++++- 1 file changed, 250 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 637358d76..d43bbac7e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -23,6 +23,10 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. ### Added +- **`--drop-xe-sessions`** ([#4852]) - a command-line verb that drops the Extended Events sessions a removed server left behind. Run it just before removing the server. `--dry-run` lists them. `--print-sql` prints guarded statements to run by hand when the server is no longer configured. +- **Collection Falling Behind self-alert** ([#4851]) - it fires when the fleet collection gate skips at least 5% of the slots due in the last hour. At least 20 slots must be skipped too. It resolves after a full hour under 1%. An hourly log line gives the slots run and skipped, how long collection waited for a gate slot, and the gate width. +- **PostgreSQL targets: a finding when pg_stat_statements keeps evicting statements** ([#4680]) - When a target's pg_stat_statements module has had to drop its least-used entries in at least 3 of the last 6 hours, analysis reports it, says Darling's statement totals can under-report, and recommends raising pg_stat_statements.max (a restart, or on RDS and Aurora a parameter-group change and a reboot). It never fires where the eviction count is unknown. +- **Per-rule plan analyzer overrides** ([#4602]) - `darling.json`'s (Darling) and `settings.json`'s (Lite) new optional `analyzer` section can disable a specific plan-analysis rule or override its severity, e.g. `"analyzer": { "rules": { "disabled": [20] } }` hides Local Variables findings. Applies to the MCP plan tools, the drill-down views, and the plan viewer window. - **Plan viewer: a Server Context card** ([#4614]) - The Plan Insights strip now opens with a Server Context card, as in PerformanceStudio's desktop viewer. It shows the server a plan belongs to: its name, edition and version, CPUs and memory, MAXDOP, cost threshold for parallelism, max server memory, and the plan's database with its compatibility level, from what the collector has stored for that server. - **Plan viewer: a Parameters card** ([#4604]) - The Plan Insights strip now includes a Parameters card, listing each statement parameter's data type, compiled value and runtime value, and flagging a runtime value that differs from its compiled value as possible parameter sniffing. - **Plan viewer minimap** ([#4596]) - The plan viewer's toolbar has a "Minimap" toggle that opens a scaled-down overview of the whole plan, with a box showing the current scroll position and a click-to-center shortcut. @@ -118,6 +122,24 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. ### Changed +- **Per-server Query Store reads fetch fewer rows** ([#4861]) - This covers the Queries grid, MCP's Query Store top read and the Trends Query Store chart. Each filtered one server's rows of the per-interval table by collection time only, which no index held. So it fetched every stored row of that server. Each read now also bounds first_execution_time, which the table's unique key holds, so older rows are dropped before they are fetched. Results do not change. On a large store, a day-long MCP read went from 3.1 s to 0.6 s. +- **SQL Server blocking and deadlock baselines count the hours collection covered** ([#4853]) - each hour-of-week's mean now divides by the days collection ran in that hour. It used to divide by the days that had events. A covered hour with no blocking or deadlocks is a measured zero, so a spike into it says the baseline measured zero instead of "first occurrence". This needs at least one captured event in the 30 days, so a server that cannot capture the events keeps no baseline. Hours with occasional events get lower means. More hours are trusted, and a trusted hour must also pass the rate multiple, so those hours alert less often. An hour whose events cluster on a few days can alert at a lower event rate than before. These baselines are now computed once per UTC day, like the CPU and I/O latency baselines. Events earlier on the current day are not part of the baseline. +- **The plan viewer's Runtime Summary lists CE model above Optimization, with Early abort under it** ([#4849]) - this matches PerformanceStudio. The Early abort label is indented when an Optimization row is shown. +- **Purge Now runs in the background and is paced like the daily purge** ([#4835]) - the command ran the whole purge before it answered, so every other command waited behind it. Its deletes were not paced, which slows collection on every server the way the daily purge did. It now answers at once with "started", or "already running" when another purge holds the slot. When the purge finishes, its totals and each raw table's result are written to the collection log under (fleet), which `get_collection_log` reads, and the viewer shows the totals. Update viewer seats along with the service: a viewer from an earlier release reads the new answer as 0 rows purged. +- **get_pg_deadlocks explains an empty answer more completely** ([#4832]) - The tool's text now names the log line prefix check beside the message language check (as does the empty answer of get_pg_deadlock_detail), and says the RDS read position is saved, so a restart resumes where the last read stopped. +- **Darling reads a quiet database's watermark from its cache instead of from the store every cycle** ([#4818]) - A database whose newest row was recent but had no new batch since the cache was seeded missed the cache on every cycle, at about 30 ms a time. +- **Query Store rows now record when each interval ended, in both apps' stores** ([#4802]) - The Query Store stats in Darling and Lite gain an `interval_end_time_utc` column next to the interval's start time. Rows collected before the upgrade leave it empty. +- **Query Store backfill: the per-tick database list no longer walks the newest chunk's index** ([#4703]) - every backfill tick listed a server's databases with a read bounded at the backfill horizon. Inside the chunk that holds the horizon that bound can only filter, so the read walked the chunk's index entry by entry, and it grew as the chunk filled. On a store monitoring 43 servers, with the horizon 20 minutes before the end of its chunk, one round of these reads cost 495,218 buffers and 560 ms (up to 24,457 buffers and 44 ms for one server). The read now starts at that chunk's start, taken from the TimescaleDB catalog, and the same round cost 1,815 buffers and 5.3 ms. On plain PostgreSQL the list comes from a walk of the index, one seek per database. A quiet database whose newest rows are in that chunk but older than the horizon is now marked done at once, so if it comes back after those rows leave raw retention, its first-contact history is not backfilled. A recorded outage gap is still backfilled. +- **Query Store views over windows longer than the raw tier now show the full window from the interval table, and say how far back it reaches** ([#4700]) - The Queries grid, the get_query_store_top MCP tool and Compose Query Store panels read the per-interval table below the raw tier's floor, down to the exact point where it provably holds the same rows raw held. The MCP tool reports history_source and effective_start, and its note says which bound set the start. +- **Analyzer wording for Nested Loops outer sides and row estimate mismatches now counts rows the way the plan viewer does** ([#4698]) - Rule 16's Nested Loops outer-side detail now prints the total actual rows when the outer input is not on a Nested Loops inner side, for example "actual 100,000 (1000x underestimate)" where it used to print "actual 12,500 (125x underestimate)" at DOP 8. Rule 5's "(N rows x M executions)" text now appears only on a Nested Loops inner side. +- **PostgreSQL targets: the top-queries views say when pg_stat_statements evicted entries during the window** ([#4680]) - Darling now records each target's pg_stat_statements eviction count (`pg_stat_statements_info.dealloc`) with its statement statistics. When the module had to drop its least-used entries in the window, `get_pg_top_queries`, the Darling Viewer and the web dashboard's Top Query Shapes say so and that the totals can under-report, and name pg_stat_statements.max as the fix; where the info view is absent the count is reported as unknown rather than zero. +- **Web dashboard: each notebook and Custom View has its own auto-refresh setting, notebooks start with it off, and slow pages back off** ([#4671]) - A labelled "Auto-refresh" control on each saved view and notebook (Off, 1, 5 or 15 minutes) is saved with the definition; notebooks default to Off and Custom Views to 5 minutes, and the fleet and server pages keep refreshing every minute. The header button now pauses all auto-refresh. A page whose last render took more than half its interval waits at least four times that render (up to 15 minutes) before refreshing again, and says when it will. +- **Query Store collection: each database's watermark comes from the store only when a cached value can't be proven current** ([#4669]) - The collector read `MAX(last_execution_time)` for every database on every cycle; on one production store that was about 6,900 s of store time and 112 million blocks read in four days. The value is now cached per server and database, advanced only after a batch commits and dropped on any fault, reconnect or backfill write, and used only while the row that set it is still inside the read's three-hour floor, so every cycle collects from exactly the same point as before, and re-read from the store at least hourly. +- **Alert pass: the forced-plan failure check reads the store only when its answer can have changed** ([#4667]) - Between Query Store collections every pass re-read two hours of one server's Query Store rows to reach the same answer; on one production store that was the second-largest statement, about 27,700 s of store time in four days. A one-row probe now decides, every Query Store write clears the saved answer, and the previous answer is reused otherwise (less any plan whose older collection has left the window), so alerts fire exactly as before. +- **Query Store collection: the orphaned per-database state cleanup runs at most once an hour per server** ([#4665]) - It ran three deletes after every Query Store cycle to catch the rare dropped database; on one production store that was about 5,600 s of store time in four days. It now runs on the first cycle after a start and then hourly, in Darling and Lite; a failed run is still retried on the next cycle. +- **The Darling Viewer's Query Store trend chart and `get_query_store_duration_trend` no longer re-scan the hourly rollup's oldest and newest hour on every load** ([#4618]) - They take the rollup's range from the coverage the service already caches, instead of a fresh `min`/`max` read over the rollup on every call. +- **Custom Views Query Store panels over a recent window read the per-interval table instead of re-sorting raw snapshots** ([#4617]) - A Query Store panel of 12 hours or more, whose window the store's per-interval table fully covers, now reads that table, with the same numbers, instead of deduplicating every raw snapshot in the window. On a large fleet that dedupe spilled gigabytes of temporary files and could run to the 60-second read timeout. +- **`get_store_metrics` reports the per-interval Query Store tables by name** ([#4616]) - Their size and growth get their own daily rows instead of being counted in "other", so a stalled purge on them shows up. - **Darling Viewer and MCP store reads can no longer spill unbounded temporary files** ([#4610]) - The viewer and mcp store roles now carry a temp_file_limit, so a Custom Views panel that would have written gigabytes of temporary files fails at once with a message naming the limit, instead of running to the 60-second statement timeout and slowing the collector's writes. - **Plan viewer minimap: resize, double-click zoom, and accuracy-coloured edges** ([#4603]) - The plan viewer's minimap panel can now be resized and remembers its size across plans, double-clicking a minimap node zooms to and selects it, and minimap edges pick up the same over/underestimate colouring as the main canvas. - **Rule 38 can now flag a Standard Edition DOP limit as a Warning** ([#4601]) - Plan analysis reads each server's collected edition and MAXDOP wherever it analyzes a stored plan (the MCP plan tools, the drill-downs, and the Darling Viewer and Lite plan viewers). A plan that ran at DOP 2 with batch-mode operators on a Standard Edition server with MAXDOP above 2 now gets a Warning; it stays Info when the edition isn't known. @@ -147,14 +169,14 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. - **The store now gives the web dashboard's and MCP server's reads up to 60 s before cancelling them, instead of 15 s** ([#4447]) - existing stores still at the shipped 15 s default move to 60 s through store schema rung V147, and an operator-set value other than 15 is left alone. Reads over 5 s are still written to the store's log. The desktop viewer's reads and the shared storage-layer readers keep their own 15 s and 30 s client limits for now. Bring-your-own stores re-run `tools/provision-roles.sql` to pick up the new value. - **The managed store checkpoints every 15 minutes instead of 5, cutting write-ahead log volume** ([#4426]) - a new marker block ships `checkpoint_timeout = 15min`, after a trial on one store kept the checkpoint sync phase far under the checkpointer's own sync-phase self-alert. - **Report alerts link to their own page: the Fleet Sweep Rollup opens the sweeps page, and the digests carry no link** ([#4421]) -- **`get_top_procedures_by_cpu` and the Top Procedures grid now route to the hourly rollup once raw ages past its retention window, instead of returning nothing** ([#4413]) - Mirrors the routing already shipped for `get_top_queries_by_cpu`; the payload discloses which tier answered and what an hourly-routed row is missing (`object_type`). +- **`get_top_procedures_by_cpu` and the Top Procedures grid now route to the hourly rollup once raw ages past its retention window, instead of returning nothing** ([#4413]) - Mirrors the routing already shipped for `get_top_queries_by_cpu`; the payload discloses which tier answered, and the columns the hourly rollup lacks (`object_type`, the handles, the I/O totals and the min/max extremes) are null, not 0. - **Sized the Linux compose store's background-worker slots and `work_mem` from the product's own - **Cache the server-scoped watermark read across a runner's lifetime, cutting repeated `MAX()` reads against `job_history`, `default_trace_events`, `system_health_events`, and `memory_pressure_events` to one seed per (server, collector) pair instead of one per collection cycle** ([#4399]) -- **Top-N-by-CPU queries now route to the hourly rollup once raw's retention has dropped the window, instead of returning empty** ([#4396]) - `get_top_queries_by_cpu` (MCP tool, Storage reader, and Viewer) degrades to the hourly continuous aggregate for a window older than raw's floor and discloses the precision loss (`tier_used`, `precision_note`) rather than silently returning nothing. +- **Top-N-by-CPU queries now route to the hourly rollup once raw's retention has dropped the window, instead of returning empty** ([#4396]) - `get_top_queries_by_cpu` (MCP tool, Storage reader, and Viewer) degrades to the hourly continuous aggregate for a window older than raw's floor and discloses the precision loss (`tier_used`, `precision_note`) rather than silently returning nothing. The raw-only filters (`min_dop`, `parallel_only`, `group_by=host_object`) keep the read on raw, and when raw holds nothing in the window the answer is status `empty` with `window_truncated` set. On the hourly tier, the columns the rollup lacks (plan hash and handle, DOP, reads, writes, physical reads, rows, spills, `avg_reads`, `is_parallel`, `distinct_texts` and the min/max CPU and elapsed times) are null, not 0, and `sql_handle` is returned. The read counts no bucket past the rollup's materialization ceiling, and `cpu_attribution` divides by the span the rollup served. - **The Performance Trends Query Store duration chart reads the interval table for windows of 48 hours or more** ([#4382]) - the table measured faster than raw from 48 hours (361 ms vs 423 ms at 48 h; 374 ms vs 1,045 ms at 7 days) and returns exactly raw's points; shorter windows stay on raw. The interval-table gate also reads raw for any window that still holds legacy rows without an interval start. - **The Darling viewer's TempDB file I/O trend now buckets like the File I/O tab's own reads** ([#4353]) - a 7-day window used to return one row per collection per tempdb file (40,324 rows measured for 4 files at 1-minute cadence); it now buckets to the chart's point budget (4,036 rows measured for the same seed) and still stamps a short window's raw points unchanged. - **Raw hypertables now re-tune their own chunk interval once a day from actual ingest** ([#4344]) - Every raw hypertable used the same fixed one-day chunk interval whatever its write rate, so a busy table's open chunk, with its indexes, could outgrow the memory meant to hold it. Darling now reads each raw hypertable's compressed ingest rate daily, alongside a RAM-derived budget, and narrows the interval one rung at a time through `set_chunk_time_interval` when the store-wide total is over budget, and widens it one rung when the total would stay under half the budget; a new rung history table records every change, alongside a per-run WAL-volume record. On by default; set `rawChunkIntervalReconcileEnabled: false` in the config file to turn it off. -- **The Queries grid and MCP's Query Store top read use a per-interval table for long windows** ([#4341]) - both used to deduplicate the whole raw Query Store slice on every read. The grid now reads a table that keeps one row per interval for windows of 12 hours or more, and the MCP read for 24 hours or more. On a 15-day seed of 3.47 million raw rows, a 7-day window took 2,754 ms against 7,065 ms for the grid, and 1,627 ms against 9,249 ms for the MCP read. Shorter windows, and any window the table's history doesn't cover yet (after an upgrade), read raw as before. +- **The Queries grid and MCP's Query Store top read use a per-interval table for long windows** ([#4341]) - both used to deduplicate the whole raw Query Store slice on every read. The grid now reads a table that keeps one row per interval for windows of 12 hours or more, and the MCP read for 24 hours or more. On a 15-day seed of 3.47 million raw rows, a 7-day window took 2,754 ms against 7,065 ms for the grid, and 1,627 ms against 9,249 ms for the MCP read. Shorter windows, and any window the table's history doesn't cover yet (after an upgrade), read raw as before. A BRIN index on the table's `collection_time`, which the service builds in the background from about 20 minutes after it starts without blocking writes (skipped below PostgreSQL 16 and on a hypertable), serves Custom Views Query Store panels of 12 hours or more where `random_page_cost` is about 1.1 or lower. - **Lite's memory clerk and File I/O trend charts bucket long windows in DuckDB** ([#4340]) - The memory clerk, File I/O latency, File I/O throughput and TempDB File I/O trend charts, and the Overview I/O lane, read one point per collection. They now gather collections into buckets sized to the window, at most 1,500 points per series. A bucket's latency is its summed stall over its summed operations, and its throughput is its summed bytes over its summed seconds. A window short enough that every bucket holds one collection plots each point at its own collection time, as before. The memory clerk picker's name list is cached for 15 minutes. - **Lite keeps one DuckDB connection open, so reads attach to it instead of reopening the database file** ([#4339]) - Every read used to open and close its own connection to the local database file. One connection now stays open for the app's life, which cut a 20-tick, ten-way Overview read on a 1 GB store from 871 ms to 43 ms. Because that connection keeps DuckDB's buffer pool resident, a periodic check trims it after a heavy read: on a 1.1 GB store, one trim took process private bytes from 1,115 MB to 148 MB. - **Lite's Performance Trends charts bucket long windows in DuckDB** ([#4338]) - The query duration, procedure duration and execution count charts read one point per collection, about 10,080 points each for a 7-day window at one collection a minute. They now gather collections into buckets sized to the window, rating each bucket as its summed work over its summed seconds, and a window short enough that every bucket holds one collection plots each point at its own collection time, as before. @@ -247,6 +269,126 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. ### Fixed +- **A shutdown during a managed PostgreSQL tool step is no longer missed** ([#4867]) - A shutdown that came just before a tool ran was missed. When the tool finished at once, the work went on to its next step. Each tool run now checks for a shutdown before it starts. Steps that stop PostgreSQL or put files back during a shutdown still run their tools. +- **Rollup reads stop at the window end** ([#4859]) - an hourly rollup read counted one extra hour when its window ended exactly on the hour. It took the whole hour that starts at the end. This affected the Queries and Procedures tabs, the top-queries and top-procedures tools, the hourly trends, one query's hourly history, and Custom Views. A Custom View on the daily rollup counted the next day when its end fell at midnight. These reads now stop at the end. +- **Grid time columns sort by time** ([#4858]) - clicking a time column's header in Lite or the Darling viewer sorts the rows by time. Before, most time columns sorted as text, so "10/1/2026" came before "9/30/2026". +- **Six Postgres time columns follow the time display mode** ([#4858]) - six time columns on the Darling viewer's Postgres tabs always showed UTC. They are Lock Stats Last Seen, Wait Sampling Last Seen, Replication Connected Since, Column Stats Captured, and Index Bloat Measured and Estimated. They now show the same server, local or UTC time as the rest of the tab. +- **Server times and charts are right across a daylight saving change** ([#4783], [#4829], [#4838], [#4840], [#4841], [#4847], [#4848], [#4855]) - Where SQL Server reports its time zone (SQL Server 2022 and later), Lite and Darling convert a server's local times with the offset in force at each time, instead of one offset for every time. Older servers keep the fixed offset. In Server time mode, Lite and the Darling viewer follow the server's time zone and read it again on every refresh. This covers chart times, picker and custom ranges and System Events (Default Trace) times in both apps, and Job History run times and the SQL Agent next run in the Darling viewer. It also covers Custom Views Default Trace markers, config-change attribution, the parameter-sensitivity checks, the long-running job alert's "Started" time, and the server-local times in Darling's MCP tools (blocking, default trace, version store, object stats and running jobs) and in Lite's (`get_blocked_process_reports`, `get_running_jobs`, `get_index_usage`, `get_pvs_stats` and `get_default_trace_events`). Lite's `get_top_queries_by_cpu` and `get_top_procedures_by_cpu` compare last run times on the server's own clock, so a server west of UTC no longer drops a query last run early in the window and a server east of UTC no longer keeps one from before it. Recommendations "Ask AI" times use the selected server's own offset at each end of the window, and the Darling viewer's "Last analyzed" line shows the selected server's time and names its zone. Lite and the Darling viewer plot each chart point at its exact instant. In the repeated autumn hour, a chart no longer folds two hours onto one, and a drill or a typed range there reads the hour it names. The spring gap shows no hour that never happened, and the Darling viewer's Overview lanes axis spans the whole window. Grid and text times in the repeated hour end in their UTC offset, so the two 01:30s read `01:30:00 -04:00` and `01:30:00 -05:00`, and time columns that show the offset sort by time. Times read from a stored server wall clock (Default Trace events, job history and the next Agent run) stay plain in that hour, because the stored time does not say which of the two passes it was. Lite's FinOps First Seen and Last Seen and its collection health times follow the display mode, not this machine's clock. A chart CSV export's time column now names its zone, such as `DateTime (UTC)`, and every row gains a 4th column, `UTC offset`, so a script that matches the header `DateTime` or expects three columns needs updating. +- **Anomaly tiles agree on the window end** ([#4854]) - every SQL Server anomaly tile now leaves out a sample stamped exactly at the window's end. This holds in Lite and Darling. Eight of the eleven tile reads used to count it and three did not. So a boundary sample counted in one tile and not in another. +- **Lite shows a refused capture read as a failure** ([#4853]) - SQL Server can refuse the read of a blocking, deadlock or long-query session. Lite used to log a refused read (error 297 or 15151) as a success with zero rows, so collection health showed the capture as working. It now shows PERMISSIONS or ERROR, as a session that failed to start does. A read refused for permission (error 297) stops that capture until Lite restarts. A read that fails with error 15151 is retried every cycle. The Capture Not Running notice now names a long-query capture that could not start. It used to call it deadlock. +- **A backward clock step no longer pauses collection or delays repeat alerts** ([#4851]) - only a clock that stepped back can put a due time more than one interval ahead. Such work now runs at once instead of waiting out the step. This covers Darling's collectors, custom alert rules, retries and periodic checks, and Lite's collectors, collection cycle and background jobs. After a step back, a repeat alert waits one cooldown from the first check that sees the step. Before, it waited the step plus the cooldown. +- **A server edited during its connect no longer runs on the old connection string** ([#4851]) - the connect checks the definition again before its first pass. If it changed, the runtime is thrown away. +- **PostgreSQL spike advice names a baseline that measured zero** ([#4850]) - an hour-of-week can see no deadlocks, or no waiting, for the whole 30-day baseline. A spike in that hour now says the baseline measured zero, and prints no multiple. It used to say "first occurrence, no baseline yet". +- **Lite's View Block Chain reads the Blocking slicer's window when a report has no event time** ([#4848]) - For a blocked process report row with no event time, the chain is rebuilt from the slicer's selected window. That selection was converted to the server's local time before it filtered the reports' event times, which are UTC, so the chain came from a window shifted by the server's UTC offset. It now reads the window the slicer selected. +- **Fleet views show WARNING when half or more of a collector's databases fail** ([#4846]) - the run still records SUCCESS. Only its note records the loss. The collector's own Collection Health tab read that note. The fleet overview, the `get_fleet_overview` MCP tool and the viewer's Overview cards did not, so they showed the collector as Healthy. On its first start after the upgrade, the service rebuilds the hourly collection-health rollup and fills eight days of it in the background. +- **`run_custom_view_panel` runs count in read latency with the web dashboard off** ([#4839]) - the run's timing went to a slot only the web dashboard filled. With the dashboard off (the default), these runs were left out. They are now recorded as composed-panel reads, the same as a panel run from the web dashboard. +- **An alert that no channel delivered is tried again a minute later instead of waiting out its cooldown** ([#4786], [#4804], [#4826], [#4827], [#4828], [#4837]) - A send fails like this when, for example, a webhook answers HTTP 429 or 5xx or times out, or the mail server cannot be reached. The alert is tried again a minute later while its condition holds, and the wait doubles after each failed try (1, 2, 4 minutes and so on) and never passes the alert's cooldown. In Darling this covers the SQL Server high CPU, blocking, blocking wait time, deadlock, poison wait, long-running query, tempdb space, volume free space, version store, file growth, long-running job, failed agent job, database state and forced plan alerts; the PostgreSQL findings, high CPU, deadlock, blocking, long-running query and poison wait alerts; and the repeating self-alerts, such as Collection Stopped, Capture Down, Agent Not Running, Collector Cost Regression and store disk pressure. Lite reports a send that no email or webhook channel received as failed, so its alerts are retried the same way, and the tray notice shows again on each retry. In both apps, "Server Unreachable" and the availability group alerts (replica disconnected, failover, data movement suspended, and in Lite, sync fell behind) are tried again while the server stays down, with re-fire off or on, and so is a Per-event deadlock, blocking or poison wait alert whose sends all failed. A delivered alert is not repeated, and neither are the Restored and Reconnected notices. A custom alert rule has no cooldown, so its wait stops doubling at 15 minutes. A severity change that no channel delivered is sent again the same way, and a resolve goes out only for an alert that was delivered. A deadlock or blocking alert that fails to send no longer uses up its count. A failed agent job is no longer recorded as announced when every channel failed, so a restart before the retry no longer loses it; a crash between a send and that record can send the same failure again. Closing Lite ends a webhook alert post that is still in flight, instead of letting it run out its timeout. +- **A removed server no longer leaves alert state behind** ([#4837]) - an alert check or send could still be running when its server was removed. It then left state behind, and the server inherited that state when it was added again. A stale availability group role could page a failover. A stale "already reported" mark could hold back the server's first alert. In Darling, a stale offline mark sent "Server Restored" on the first connect. A check that finished after the removal could also claim an availability group for the removed server, so the group's other monitored replicas stopped judging it. The removal now drops that state, and a check or alert send that finishes afterwards records nothing. +- **The daily retention purge paces its deletes** ([#4835]) - it deleted rows at full speed, and the store wrote about 6 GB of WAL in 7 minutes. The PostgreSQL checkpoint running at the time took 23.5 seconds to sync, and collection slowed on every server. The purge now waits between batches to hold the store's WAL to half of what its checkpoint settings absorb. The collection log shows the store's WAL during the purge and how long it waited. +- **The Store Checkpointer Pressure alert sees one long checkpoint sync** ([#4835]) - it judged only the hour's average sync per checkpoint. A single 23-second sync among several short ones stayed under the 10-second bar. It now also judges the longest single sync, sampled once a minute. +- **get_store_metrics reports the longest checkpoint sync the alert judges** ([#4835]) - the hour's longest single sync is stored with the hourly checkpointer counters. The tool shows it and judges it against the same 10-second bar, so the tool and the alert agree. +- **Query Store trends no longer read low after a quiet interval** ([#4833]) - Query Store stores no row for an interval with no executions, so the next interval's rate was divided by the gap plus its own length. It now uses the interval's own length, for rows collected after the upgrade. The `get_query_store_duration_trend` description in both apps now describes that rate: its head says only a point with no stored end and no earlier point has null rates, and its guide text says each point is rated over its own stored interval. +- **Editing a server's host in the Darling viewer keeps its favorite star with the server** ([#4831]) - Before, the star stayed behind under the old address and marked any server added there later, and unchecking it in the same edit did not clear it. +- **get_store_metrics no longer reads two days of store growth as one** ([#4830]) - A day whose last snapshot came early in the day, after the service was down, was compared with the next day's last snapshot almost two days later and shown as one day's growth. Such a pair is now left out, and the tool says its times are UTC. +- **Darling and Lite count the I/O anomaly tiles' samples the same way** ([#4820]) - Darling counted a row once for both the read and the write tile, so it counted some read-latency tiles that Lite correctly left out, and the sample count on those tiles differed between the products. Darling now counts read samples and write samples separately, as Lite does. +- **A tile's peak time no longer changes between runs when two samples tie** ([#4820]) - when two samples tied on the peak value, the time shown for the peak was whichever the engine reached first, so it could differ from one run to the next with no change in the data. Both products now show the later of the tied samples, for Lite's CPU tile and for every Darling tile that reports a peak time. A Darling tile's peak time also skips a sample with no value, as the tile's peak value already did. +- **Lite's archive views stay readable while compaction swaps files** ([#4816]) - A query that ran during a compaction, or while a run finished the swap an interrupted run left behind, could find an archive table empty or missing rows for a moment. +- **Lite no longer counts archived rows twice during an archive run** ([#4816]) - a new archive file was visible before its rows were removed from the database, so a query in between counted them twice. The file is now added and the rows removed under one lock. +- **"Recurring at this hour" survives a daylight saving change** ([#4814]) - After the clocks changed, older findings shifted by an hour, so the label went missing for up to three weeks, and a job that had not moved could read as "Maintenance window moved". +- **A nightly job no longer reads as "Maintenance window moved"** ([#4814]) - Rows from other weekdays were compared as if they were the same weekly slot. +- **Reading a PostgreSQL log no longer records a false error when the read starts inside a multi-byte character** ([#4813]) - a first read, or a read after a fallback, starts 4 MB before the end of the log, and when that byte was inside a multi-byte character PostgreSQL refused the whole slice and the cycle logged an error that blamed a planted byte. The read is now repeated in the same cycle from 1, 2 and 3 bytes later, and the error row is written only if all four attempts are refused, with a message that says a split character is possible and that later reads resume from a full line. +- **A byte that is invalid in the database encoding costs one log read per cycle, not four** ([#4813]) - the retry from a later start cannot cure an invalid byte anywhere in the 4 MB window, and it is refused at every start on every cycle. After a cycle that used every start and was refused at each, the collector now makes one read per cycle until a read succeeds, and goes back to retrying after that. The memory is kept in process, so a restart pays one four-read cycle. +- **PostgreSQL error rows under the pgBadger log prefix now carry their user and database** ([#4813]) - a `log_line_prefix` written as `user=%u,db=%d` gave a null user and database on every error row, because only `user@db` was read. Both `user=` and `db=` are read now, wherever they sit in the prefix. +- **The plan-capture readiness read now reports a log prefix the log readers cannot read** ([#4813]) - a `log_line_prefix` with an unbracketed process id (`%m %p`, `%t %p:`, `%t [%p-%l]`) makes the stderr log readers drop every line, and the target looked quiet. A new `log_line_prefix_readable` row says whether the prefix starts with a timestamp and carries `[%p]`. +- **A deadlock report cut at the end of a read waits for its last lines instead of being stored as a fragment** ([#4813]) - a report whose HINT line the read did not reach was stored from the lines the read had, and one cut inside its DETAIL block was stored again under another hash when the next read had it whole; on the RDS log API, which never offers the rest again, only the fragment was stored. The newest report of a self-hosted read with no line after its DETAIL block is now left for the next read, which covers those lines again. On RDS the report is held until the next chunk and stored as it is if it is still unfinished after one more chunk, if its file has ended, or if it is larger than 1 MB. A report that other log lines follow is stored at once. +- **The fleet sweep bands a server outage as No Data instead of Warning** ([#4811]) - While a server could not be reached, its "Collection Stopped" alerts counted as data. The outage then showed as a Warning with an alert count, counted as reported, and hid the stale-collection item. +- **A collector that fails on half or more of its databases now bands Warning** ([#4811]) - A run that skipped half or more of its databases still recorded success, so Collection Health read Healthy with no errors. +- **The Darling viewer's Collection Health tab no longer re-reads the store's schema version on every refresh** ([#4811]) +- **Plan-forcing advice says when SQL Server's own recommendation names the proposed plan as the regressed one** ([#4810]) - The plan-force bot's journal now records a blocked decision for that plan instead of a would-force one. Each target in the script now also shows how long ago its plan was last seen. +- **Lite no longer archives the same rows twice after it is stopped in the middle of an archive run** ([#4808]) - If Lite was closed or killed after it wrote an archive file but before it deleted those rows, the next run archived them again, and the archive counted them twice. Leftover partial files from an interrupted run are now removed, and the archive views are rebuilt even when compaction fails. +- **Lite's Query Store repair also covers archive months that compaction split into parts** ([#4808]) - Rows in those part files were never repaired. +- **Darling retries a store that is at its connection limit when the service starts** ([#4806]) - Before, a store that was briefly full at start, after a restart storm or while another application held its connections, stopped collection until someone restarted the service. Other resource errors (disk full, out of memory, a configuration limit) still stop it, because they need an operator. +- **The compose `darling` container reports unhealthy when a startup failure has stopped collection** ([#4806]) - It used to show as running although it never collected. +- **The docs now name the oldest supported builds: SQL Server 2016 SP2 and SQL Server 2017 CU3** ([#4805]) - The query stats and procedure stats collectors read columns that older builds do not have. The docs said SQL Server 2016 and named no service pack. +- **A webhook that never answers no longer holds up alert delivery** ([#4803]) - each webhook post (Teams, Slack, generic and PagerDuty, and the test sends for them) waited up to 30 seconds with no way to cancel it, and a delivery makes those posts one after another. Each post now gives up after 10 seconds and records the channel as failed with "timed out after 10 seconds" as the reason. When the Darling service stops, an alert post in flight is cancelled instead of waited out, and the abandoned delivery writes no history row. Analysis-finding alerts, and every alert in Lite, get the 10-second bound but are not cancelled. +- **Query Store backfill now logs a warning when it cannot read its list of work** ([#4801]) - when the backfill's read of the databases to work on, or of a database's stored floor, failed, both apps logged it only at Debug and stopped filling tails and outage holes with no visible sign. Both apps now log one warning at the first failure of a run of failures, and again on the next failure after a read succeeds. +- **Query Store backfill keeps a slice size that works instead of swinging back to the size that timed out** ([#4801]) - after a timeout the slice halved, but any completed slice sent the server back to the full hour, so a database whose 30-minute slice fit and whose 60-minute slice did not wasted a 60-second read on every other tick. The slice still halves after a timeout, now keeps its size after a slice completes, and widens by one step (never above one hour) after three completed slices in a row. Both apps. +- **PostgreSQL analysis: a configuration setting shown under another finding now carries its own advice** ([#4799]) - A setting that hung off another finding, such as maintenance_work_mem under an autovacuum backlog, was described as "see its card", and only analyze_server shows that card. The viewer, get_analysis_findings and the e-mail showed neither the setting's advice nor its fix. The finding's text now names each such setting with its own headline and adds its fix, so every place that shows the finding shows the advice. +- **The advice for turning autovacuum back on now names the provider's parameter group** ([#4799]) - The advice said only "autovacuum = on in postgresql.conf", which managed PostgreSQL (RDS, Aurora, Azure, Cloud SQL) does not let you edit. It now names the provider's parameter group or server parameters, which apply the change themselves, and keeps the reload for postgresql.conf. +- **A webhook channel that keeps failing no longer goes unnoticed while another channel delivers** ([#4798]) - A Teams, Slack, generic or PagerDuty channel could fail for weeks with nothing but log lines while another channel delivered every alert. A channel that fails three times in a row now raises a "Notification Channel Failing" alert in Darling and a tray notice in Lite, and clears with "Notification Channel Recovered" when it delivers again. Turning a failing channel off closes the alert with an entry that says the channel was turned off. The alert and the notice name the channel and the count and carry no error text, since a webhook error can carry the webhook URL. +- **The configuration change card no longer calls a metric "resolved" when only minutes have passed since the change** ([#4797]) - Until the after-window reaches an hour, a metric that appeared or vanished is shown as not yet comparable. +- **A server that is down when the Darling service starts now gets a "Collection Stopped" alert once the threshold passes** ([#4796]) - the check was skipped until the service had seen the server online, so a server that stayed down across a service restart got no "Collection Stopped" alert, and no reminders, until it came back. Staleness is now judged from the later of the server's last success and the service start, so the alert fires once the staleness window (30 minutes by default) has passed since the start, and its text counts the minutes from the start. A healthy server whose last success predates the restart still does not alert at startup, and a server that reconnects after a restart is judged from the start rather than from its pre-restart rows. +- **Adding a server whose id matches a different server's no longer overwrites it in the Darling viewer, and Lite now refuses to add or edit onto it** ([#4794]) - Two different servers can hash to the same id. In the Darling viewer, Add Server and Add Multiple Servers wrote with an upsert that rewrote the server already holding it: that server stopped being monitored and its history showed under the new one, with no message. Add Server now refuses the add, saves nothing and names the server that holds the id, and adding a server that is already monitored at the same address shows the "already monitored" message instead of rewriting its row. Add Multiple Servers refuses that row, counts it apart from added, skipped and failed, and shows the count in its status text. The `--add-server` totals line now counts collided servers, and the command exits 1 when a server was not added because its id collides. In Lite, both servers were added and then collected into one history, so each tab showed both servers' rows with no message. Lite's Add Server and Edit Server now refuse a server whose id a different server holds and name that server, and adding the same server twice shows the "already monitored" message. Lite's Add Multiple Servers marks such a row "Not added", counts it apart from added, skipped and failed, and the main window's summary shows the count. An edit that keeps its id is still allowed, so a pair saved earlier stays editable. Importing follows the same rule: Lite's Import Settings no longer adds a server whose id a different server holds, and the Darling viewer's migration and Import Settings no longer leave such a server out silently as if it were already monitored. Each refuses it and reports it: Lite's Import Settings message shows the count and its log names the entry and the server that holds the id, and the Darling viewer logs both servers, with its Import Settings message showing how many servers it did not add. +- **`get_store_metrics` no longer reports a gap of several days in the store's size history as one day's growth, and it marks today's partial day as partial** ([#4792]) - The whole-store `daily_growth` labeled the difference between two neighboring daily points with the later day without checking that they were a day apart, so the first day after a stretch with no snapshots carried the whole stretch's growth as a single-day spike, and the per-server ingest rate built on it overstated what onboarding a server costs. Today's point, the latest snapshot of a day still in progress, was shown as a full day. A day is now compared only with the calendar day before it, and a wider pair is left out of the list (a hole in `daily_growth` is missing snapshots, not zero growth). Each point also carries `metric_time` (the snapshot it was read at), `span_days` (always 1) and `partial` (true for today), which makes each point about 74 bytes larger. +- **The store checkpointer alert no longer tells you to raise max_wal_size for every requested checkpoint** ([#4791]) - The informational Store Checkpointer Pressure alert, and the checkpointer note in the MCP store metrics, called any requested checkpoint "WAL-forced" and said the store had written more WAL than max_wal_size allows. PostgreSQL also counts a base backup or a CHECKPOINT statement as requested, and its counters do not say which one happened, so the suggested remedy was sometimes wrong. The alert's text now says "requested checkpoint(s)". The alert detail and the note now list all three causes, and when a requested checkpoint is what fired them, they also point at the log_checkpoints lines that say which and suggest raising max_wal_size only for the WAL cause. The get_store_metrics tool description says the same about its requested field. The alert fires on the same conditions and stays informational. +- **`mute_analysis_finding`, in Darling and in Lite, no longer mutes the pattern on whichever matching server sorts first** ([#4790]) - a `server_name` that matched more than one server (a partial name, or a display name shared by per-database and read-only registrations of one machine) muted the pattern on whichever sorted first and echoed the caller's spelling. Both products now resolve it to exactly one server, Darling the way `remove_server` does: a name that matches more than one server answers `ambiguous` with the candidates and mutes nothing, a name that matches nothing (a blank one included) answers `not_found`, and a successful call reports the resolved server name. `remove_server` resolves the same way: a name that matches one registration's own server name exactly, in the same letter case, picks it even when a sibling shows that name as its display name, while a partial name, a display name or a different letter case that matches several registrations still answers `ambiguous`, and the successful answers of both tools now say which kind of registration was picked (`kind`: plain, read-only or per-database). +- **Repeating a `create_mute_rule` call no longer creates a second identical rule** ([#4790]) - a client retry, from MCP or the web route (`POST /api/mute-rules`), left two identical mute rules. A create that repeats an enabled, unexpired rule's scope, patterns and expiry now inserts nothing and answers `already_exists` with the existing rule's id (HTTP 409 on the web route). A disabled or expired rule does not count, and a different expiry is a different rule. +- **`add_servers` reports every server in the batch when a later entry fails to save, and no longer says "added" for a save that wrote nothing** ([#4787]) - `add_servers` now reports every server in the batch when a later entry fails to save, instead of reporting the whole call as failed after it had already added the earlier ones: the failed entry and the ones after it come back as the new status `not_saved` (counted as failed) and the earlier ones stay `added`. It rejects a literal SQL password or service principal client secret on Linux, where it cannot be encrypted, and points at the `env:` or `file:` reference form. When the save wrote nothing because another server holds the same id, it answers `collides` (or `duplicate` for a concurrent add of the same server) instead of "added". +- **The plan-force bot no longer reads a database missing from the newest automatic-tuning capture as "automatic correction off"** ([#4784]) - it uses that database's last state from the past 24 hours, and blocks the database as unknown when there is none. +- **After a failed state read, the plan-force bot looks at the same plans again after 1 hour instead of holding them for a day** ([#4784]) - the blocked row from a failed read now holds its plan for 1 hour, and every other blocked row keeps the 24-hour cooldown. +- **Web viewer: a page left loading in a hidden or paused tab refreshes normally again** ([#4781]) - A page whose data finished loading while its browser tab was hidden, or while auto-refresh was paused, used to count the whole hidden or paused time as its load time. It then waited 15 minutes before its next refresh and showed "(slow page)". The page now counts its load as finished when its data arrives, and refreshes on its normal schedule. +- **The PostgreSQL plan-regression finding now reports the worst regression when more than 50 statements changed plans in the window** ([#4780]) - the plan-change read kept the 50 lowest query ids, so the worst regression could be dropped, or no regression reported at all when the kept rows had too few calls. The read now keeps the statements the pick would choose from, worst first. +- **A standby that was removed or replaced no longer outranks the live standbys for the rest of the window** ([#4780]) - its last row stayed in the window, so the finding described it in the present tense, kept its synchronous weighting and reported the whole table's span. Live standbys now rank first, and a standby that stopped reporting gets advice that says when it was last seen. +- **A dropped replication slot no longer grades Critical for the rest of the window** ([#4780]) - an inactive, growing slot kept grading Critical from its last row after it was dropped, and its inactive hours kept growing. A slot or database that stopped reporting now grades nothing, ranks behind the ones still reporting, and measures its inactive hours to its last row. +- **The alert history now records whether each notification channel delivered or failed** ([#4779]) - an alert that reached one channel and failed on another was stored as a plain delivery, so a broken channel was invisible while a sibling kept delivering. The route record on each history row now says `delivered`, `failed` or `not attempted` per channel, in both Darling and Lite, and `get_alert_history` returns it as `outcome`. No error text is stored, because a webhook error can name its endpoint URL; `send_error` is unchanged. +- **The alert notebook finds an alert's resolution even when it came late, was dismissed, or sat behind many newer alerts** ([#4778]) - The notebook's status only looked at the first 24 hours after the alert, at most 200 rows, and skipped dismissed rows. A store alert that cleared more than 24 hours later, a Cleared row hidden by "Dismiss all", or a fleet-level alert whose resolution sat behind 200 newer alerts from other servers read "Fired again" or "Unknown" instead of "Resolved at ...". The status now reads the first resolution row and the first later firing of the metric directly, with no 24-hour limit and no row cap, dismissed rows included, and scoped to the alert's server when it has one. +- **A fleet-level store alert with no resolution reads "No resolution recorded" instead of saying collection stopped** ([#4778]) - A store self-alert with no resolution row and no later firing read "Unknown (not collected since ...)" in the alert notebook, which says a collector stopped. Store alerts have no collector, and the process that evaluates them serves the page, so it now reads "No resolution recorded". +- **Darling now sends alert email through notification routes when the default recipient list is blank, instead of treating email as not set up** ([#4777]) - A notification route that named email recipients as its only destination sent nothing when the SMTP settings had a server and a from address but no default recipients, because email counted as set up only with the default list. Email is now set up with a server and a from address, and a route's recipients are enough to send. An alert that no route covers and that has no default recipients is skipped without a failed row or a cooldown, and an alert nothing else delivered is recorded as undelivered. The Send Test Email button still needs the default list. The Settings window's Validate message for a blank default list now says what the list is for. +- **`--configure-network` keeps the web listener's `tls` and `oidc` settings** ([#4775]) - Re-running the wizard for the web listener no longer drops `web.network.tls`, `web.network.oidc` or any other setting it does not ask about, so the dashboard keeps HTTPS and OIDC sign-in after the restart. The MCP and store blocks keep their other settings too, and the wizard prints a `Kept ... from the existing darling.json.` line naming what it kept. +- **A repaired hourly hole no longer leaves a partial day in the daily rollups, and days an earlier repair left short are rebuilt once after the upgrade** ([#4739], [#4763]) - After an hourly hole repair, the daily rollups that already held a bucket for the repaired day were not refreshed again, so a day older than the daily policy's 3-day window stayed partial and became the only copy once the hourly aged out. The repair now force-refreshes those complete days (at most 3 per daily per range), including the day-grain daily behind the interval daily, and the seam repair log line reports the holes still standing. After the upgrade, the service also checks each daily rollup against the hourly data under it, once, and rebuilds any day an earlier hourly repair left short, one day at a time. The repair logs each daily day it rebuilds with the seconds it took, and a day rebuilt after a repaired seam range is no longer rebuilt twice and counted twice in the repair's chained-days count. +- **Darling upgrade no longer says "New build in place." after copying nothing** ([#4762]) - when the staging folder's name contained [ or ], the upgrade script's copy matched nothing and threw nothing, so it reported success over the old build. It now copies the folder by literal path, and stops with a plain message if the installed service executable does not match the new build's after the copy. The installer no longer skips hardening older darling.json backups, or looks up its own files by wildcard, in an install folder with [ or ] in its name. +- **An older Lite refuses to open a data file from a newer version** ([#4753]) - An older Lite pointed at a data folder that a newer Lite had upgraded kept running, and every batch for the 13 tables that gained columns in schema versions 60 to 65 failed with only a log line to show for it. Lite now stops at startup with a message that names the file and both schema versions, and does not write to the file. +- **Lite re-applies its newer columns on every start** ([#4753]) - A data file where one of the columns added in schema versions 60 to 65 failed to apply, or was stamped without it, kept failing that table's collection on every start. Lite now adds any of those columns that are missing on each start, so the file gets it back. A column that still cannot be added is logged as an error naming the table and column, and the next start tries again. +- **After the computer sleeps and resumes, Lite runs each due collector once** ([#4753]) - After a sleep, each collector that was due ran twice about a minute apart, which wasted queries and stored an extra sample per collector. Lite now starts the cycle at the current collection slot after a resume, so each due collector runs once. +- **Two analyze_server calls that overlap no longer give the second one a false all-clear** ([#4742]) - The Darling and Lite MCP servers shared one analysis service, so a second analyze_server call made while the first was still running got "No significant findings" for a server it never analyzed. Each call now runs its own analysis in both products. +- **In Lite, a server that just went down no longer holds a connection slot while its retries wait** ([#4740]) - Each connect attempt now takes a slot and gives it back when the attempt ends, so a server that went down since its last status check no longer holds a slot through four connect timeouts and seven seconds of backoff, and can't stall collection for the other servers that way. +- **The self-hosted PostgreSQL log tail no longer stays on the outgoing log file after a rotation in the same second** ([#4738]) - When a size or age rotation put the old file's last write and the new file's first write in the same second, the tail could pick the old file as current on the stderr, csvlog or jsonlog route, so lines in the new file were not collected until the next line was logged in a later second; it now breaks that tie by file name, so the newer file is read. +- **Store upgrade: copy mode is chosen from the data folder's own volume and a finished size count** ([#4725]) - The pre-upgrade headroom check read the free space of the data directory's drive letter and took a size walk an error had cut short as the size, so a data directory on a folder-mounted volume, or one whose walk failed partway, could be upgraded in copy mode and fill the disk (the store came back on its old major, but was down for the attempt). The free space is now read for the data directory's parent itself, and an unfinished size walk falls back the way too little room does: hard-link mode where the volume supports it, otherwise no upgrade. +- **Store upgrade: hard-link mode removes the pre-upgrade directory even when a settings carry fails, after saving the old `postgresql.conf`** ([#4725]) - In hard-link mode the moved-aside pre-upgrade directory shares its files with the upgraded cluster and is not a rollback copy, but a failure carrying `postgresql.auto.conf` or the operator's `postgresql.conf` lines left it in place for two service starts. It is now removed as soon as the carries are done with it, whether they succeeded or failed. The old `postgresql.conf` holds the operator's lines below the `darling-managed.conf` include and exists nowhere else, so every major upgrade now first copies it beside the data folder as `postgresql.conf.pre-upgrade` (restricted to the service account like `postgresql.auto.conf.pre-upgrade`, covered by `--harden-files`, and kept until a later major upgrade replaces it), and hard-link mode removes the directory only once that copy exists, or when there was no `postgresql.conf` to copy. If the copy cannot be made, the directory is kept with a warning and ages out after two service starts, as before. The warnings for an operator line the new major rejects now name where the original line is kept (the saved copy) instead of the pre-upgrade directory. +- **Store settings migration: a repeated attempt restores the newest backup of `postgresql.conf`** ([#4725]) - A failed migration attempt restored the FIRST attempt's backup, so edits made to `postgresql.conf` between attempts were lost. A new backup is taken when the file changed since the last one, and a failed or resumed attempt restores the newest. +- **Store free-space checks and reports read the data folder's own volume** ([#4725]) - The free-space check before `--backfill-rollups` and the plan dimension compaction, the WAL sizing (`max_wal_size` and `min_wal_size` at every service start), the store host profile (the start-up log, `--check-settings` and the MCP host tool), and the store disk-pressure alert read the size and free space of the data directory's drive letter. A data directory on a volume mounted at a folder was therefore judged by another volume: a backfill or a compaction could start on a volume with no room for it, the WAL ceiling and the reported size and free space came from the wrong disk, and the disk-pressure alert watched a disk the store is not on. On Windows they now ask for the data directory itself, as the upgrade check does. The alert's free figure is now the space available to the service account, as the other reads use, which differs from the volume's total free space only where the account has a disk quota. Where that read fails, each keeps its old answer for an unreadable volume (the verbs refuse, the WAL sizing is skipped and the WAL settings in force stay, the host profile shows a volume that is not ready, the disk-pressure alert gives no signal that check). The filesystem name in the host profile still comes from the drive letter, and other platforms keep the old read. +- **Lite archive compaction no longer loses a month when a promote fails** ([#4718]) - hourly compaction deleted the month's existing parquet file before moving the merged replacement in; a move that failed (a scanner or backup agent holding the fresh file) then deleted the merged copy too and the month was gone, with part files losing both old parts and double-counting the new rows. The swap now sets the existing files aside, moves every merged file in, and deletes inputs only after all of them are in place; a failed or interrupted swap is undone or finished on the next run. +- **The size-triggered Lite reset no longer deletes the database after a failed export** ([#4718]) - when the store hit its size limit and one table's export failed, the reset ran anyway and that table's hot window was lost; a kill mid-export could also leave a truncated file that hid every archived month of the table. Exports now go to temp files, nothing is promoted until every export succeeded, a failure leaves the database and archive untouched and backs off, and a kill between export and reset no longer duplicates the hot window. +- **The tab Refresh button no longer runs collectors during a Lite database reset** ([#4718]) - a reset could delete the database under a Refresh run and lose everything collected until the next restart; the run now registers with the reset gate like the scheduled and tab-open collections. +- **Lite archive views now include compaction's part files** ([#4718]) - a month that compaction split into `_ptNNN` files dropped out of the archive views, and a table with only part files showed no archive at all. +- **Importing a previous Lite install no longer overwrites this install's part files, and imported files now expire with retention** ([#4718]) - an imported part file of the same month and table was compacted over the local one; imported files also never aged out because their prefix hid the date. +- **A failed archive export no longer leaves a partial file on disk** ([#4718]) - When DuckDB failed partway through writing an archive export (out of memory, disk full), the partial `.tmp` file stayed behind, and the size-triggered reset repeated that every 15 minutes. The reset and the hourly archival now delete the partial file when the export fails. +- **Reads of an archive view no longer fail after old archive files expire** ([#4718]) - Expiring a table's last multi-part month (or its last single-file month while part files remained) made every read of that table's history view fail until the next hourly refresh. The views are now rebuilt right after retention deletes files, and after the cleanup of files a killed size-triggered reset left behind. +- **A managed store moved aside by a failed upgrade is no longer replaced by an empty one** ([#4717]) - When an in-place PostgreSQL upgrade could not move the data directory back, the next service start initialized an empty store in its place, wrote a new superuser credential over the old one, and deleted the real store as an expired rollback copy two starts later. A start now refuses to initialize while a `-old-*` or `-upgrade-*` sibling holds a cluster and names the rename to do by hand, and the rollback sweep counts no start until a cluster is back at the data directory. +- **A start after an interrupted store upgrade resumes it instead of refusing every start** ([#4717]) - When the service died between the runtime swap and the upgrade's commit, or the revert was refused, every later start stopped with "the PostgreSQL 17 binaries are not on this host" although they were in `pg-runtime-prev`, and deleting the runtime stamp to force a retry deleted them. The start now finds the store's binaries there, the runtime update never empties `pg-runtime-prev` while the store still needs it, and the refusal's hint names a path a start reads. +- **A failed store upgrade that left no startable store stops the start** ([#4717]) - A failure that could not put the data directory back, or could not revert the runtime, used to let the start go on, into the empty-store initialize above or onto PostgreSQL 18 binaries in front of a PostgreSQL 17 cluster. It now stops with the hand step in the message, and the worker does not retry it. +- **The pre-upgrade rollback copy survives the two service starts it is kept for** ([#4717]) - A start the worker retried in the same process counted as two starts against the copy. The sweep now runs once per service start. +- **PostgreSQL 13 servers read through csvlog are no longer reported as quiet** ([#4714]) - The csvlog parser required exactly 26 columns, but PostgreSQL 13 writes 24 (`leader_pid` and `query_id` arrived in 14), so every PostgreSQL 13 record was rejected. Log events, deadlocks and plan capture came back empty, and the run counted one discard however many records were rejected. The parser now accepts 24 or 26 fields, and a read it cannot use reports how many records it dropped. A PostgreSQL 13 auto_explain capture has no query id, so it is stored as an orphan plan (query id 0) instead of being counted as forged. +- **RDS and Aurora PostgreSQL log reads now collect the lines a log file received before RDS rotated it, and resume where they stopped after a restart** ([#4713]) - Deadlock, log-event and plan-capture reads opened only the newest RDS log file, so whatever the previous file received after the last read was never collected, a new stderr file was read from its last 10,000 lines instead of its first line, and the read position lived only in memory, so a restart re-read just the last 10,000 lines of the newest file. A read now finishes the file it was reading before it moves to the newest one, in the same cycle (at most four reads per cycle), and a new file is read from its first line. Each read saves its position to collect.collector_state once its rows are stored and loads it at start; a saved file RDS no longer lists falls back to the newest file's last 10,000 lines. Files that rotate away entirely between two reads are still not read, but the run now records how many in log_files_skipped_by_rotation, and a missing saved file in log_resume_file_missing, with a sentence on the collection_log row. The read also lists every log file even when the listing is paged (up to 20 pages), so a listing too long for one page can't hide the newest file or the file a saved position names; past 20 pages it uses the files it has and logs a warning, at most once an hour for each instance. +- **Down servers no longer use up the fleet's collection slots, and a command timeout no longer drops a healthy connection** ([#4712]) - A connect attempt against a dead server takes about 15 seconds to fail, and it ran inside the fleet collection gate, so about 22 down servers filled the default gate and healthy servers queued behind them. Connect attempts now run in their own small gate, and the retry delay grows from 60 seconds to about 4 minutes (with jitter) while a server stays down; it resets when the server connects or its definition is edited. A definition edit made while a connect attempt is running is no longer lost: the attempt remembers the definition it used, and a stale attempt installs nothing and lets the next sweep connect with the new definition. Separately, a SQL Server command timeout dropped the server's connection and re-ran the Extended Events setup and on-load snapshots even though the connection was still good, costing a stressed server 75 seconds or more of collection; the service now drops the connection only when a probe on a fresh connection fails. +- **A memory grant that used none of its memory now gets an Excessive Memory Grant finding** ([#4705]) - Rule 9 skipped any plan whose `MaxUsedMemory` was 0, so a grant of 1 GB or more that the query never used, the largest possible waste, got no finding. The parser now records whether the plan carried a `MaxUsedMemory` attribute, and the rule fires for a grant of at least 1 GB that used 0 KB with the message "Granted N MB but the query used none of it." A plan with no `MaxUsedMemory` attribute (an estimated plan, or one with no runtime grant info) still does not fire, and the ratio message and 10x threshold for a grant that used some of its memory are unchanged. +- **Estimated Plan CE Guess names the predicate behind each default guess and follows the plan's CE model version** ([#4705]) - Rule 33 called the 30% guess an equality guess (it is the inequality guess), called 10% and 1% inequality guesses that no predicate produces, never looked for the real equality guess, and ignored the statement's CE model version. It now labels each guess for what produces it, detects the equality guess (the square root of the row count from CE 120, the row count to the power 0.75 under CE 70), and reports a band only when the plan's estimator has it. The warning type, the 100,000-row minimum and the rest of the message are unchanged; the label text changes. +- **Expensive Operator no longer names an exchange operator** ([#4705]) - An exchange (Gather, Distribute or Repartition Streams) spends most of its elapsed time waiting on the operators around it, so Rule 35 could name the exchange while the operator next to it did the work. Rule 35 now skips exchange operators; the operator that did the work, such as a spilling Sort, keeps its warning. The 20% share, the 1,000 ms floor, the warning type and the message text are unchanged. +- **Self-hosted PostgreSQL targets no longer lose the log lines written just before a log rotation** ([#4704]) - The log tail keeps a resume marker per target, finishes the previous log file before reading the newest, and says so on the collection log when a file is gone or the log outgrew the read window. When one read holds more deadlock reports or captured plans than the collector's row limit, the newest are now kept, and the collection log says the limit was hit. +- **A Query Store backfill database that keeps failing no longer blocks the databases after it on the same server** ([#4703]) - The backfill runs at most one slice per server per pass in a fixed database order, so a database whose slice always failed stayed first in line and the databases behind it never got a slice. In Darling and Lite, a database that fails 3 slices in a row is now skipped while any other database on that server has work, one Warning is logged when it starts being skipped, and it is retried only when no other database has work, longest-ago failure first. A completed slice clears its count. The slice window still narrows per server after failed slices, as before. +- **The plan viewer's row label, edge colours and minimap, and analyzer rules 5 and 26, no longer misread operators in a parallel zone** ([#4632], [#4688], [#4698]) - Actual rows in a plan are a total, the estimate is per execution, and the execution count is a real loop count only on the inner side of a Nested Loops join (elsewhere in a parallel zone it is the thread count). The node label now prints the totals against the expected rows, for example "609 of 2,983 (20%)" (on a Nested Loops inner side the expected rows are the estimate times the executions, so "609 of 613 (99%)"), with just enough decimals that the numbers agree with the percentage: a Key Lookup that ran 117 times for 1 row reads "1 of 1.128 (89%)". An operator that returned exactly its estimate used to read "12 of 100 (12%)" at DOP 8 and, from DOP 11, was drawn in the critical colour; it now reads "100 of 100 (100%)" and stays neutral. The edge colours and the minimap follow the same rule: accurate operators in a parallel zone at DOP 11 or more stop drawing Blue, LightBlue or FluoBlue edges, and real misestimates in a parallel zone now show their full tier (a 20x underestimate at DOP 11 draws LightOrange, where it drew no colour before). On a Nested Loops inner side, the edge colour comes from the same ratio as the node's label, so an operator that ran thousands of times no longer gets a worst-misestimate red edge while its label shows the estimate was close. Rule 5 no longer reports a false overestimate on a parallel operator (a DOP 8 join with estimate 2,983 and 609 actual rows used to read "Estimated 2,983 vs Actual 609 (76 rows x 8 executions) - 39x overestimated" and is now silent), and rule 26 now reports "Row goal active: estimate reduced from 1,000 to 100" for a parallel scan that returned 500 rows, where it used to be swallowed. +- **Selecting the plan viewer's "+" tab through UI Automation (screen readers) adds one new sub-tab, not two** ([#4688]) - a UI Automation select on "+" sets focus on the tab and then selects it, and each counted as a click on "+", so one select added two empty "New Plan" sub-tabs in Lite and in the Darling Viewer. The plan viewer now treats the pair as one gesture and re-selects the sub-tab it just added. A mouse click or keyboard selection of "+" still adds one tab, and two separate selections still add two. +- **The Darling Viewer's server list is no longer a solid white box when the store is unavailable** ([#4685]) - While the store is unavailable the sidebar server list is disabled, and the stock list template painted it white in every theme, out of place next to the themed sidebar. Each theme now has its own list template: a disabled list keeps its themed background and border and dims instead. The same applies to the tag list in Manage Tags, which is disabled on a read-only seat. Enabled lists render as before. +- **Custom tab headers and the status-bar collector text now use readable, theme-following ink** ([#4682]) - In Lite, a selected server tab or Plan Viewer tab, and in Lite and the Darling Viewer a selected plan sub-tab, drew its label in the page text colour on the accent fill (1.99:1 on Dark, 2.71:1 on Cool Breeze); it now uses the theme's on-accent ink. In both apps, the status bar's "Collectors: N OK / erroring" text kept the previous theme's colour after a live theme switch until the next 30 second tick; it now follows the switch. In Lite, "Logging: BROKEN" uses the theme's critical text colour instead of a fixed red that was under 4.5:1 on every theme. +- **Web dashboard: a slow notebook or Custom View panel is no longer re-run while its previous read is still running** ([#4671]) - Composed panels read through a request the page's in-flight guard did not count, so a periodic refresh could fire a panel again while its last read was still running on the store. Panel reads are now counted like every other page read, so the page waits for them and the new slow-page back-off measures them. +- **Custom Views: window-total Query Store measures are no longer labelled "(ratio)"** ([#4664]) - "Query Store total CPU time" and "Query Store total duration" showed "(ratio)" in the measure pickers although they divide by nothing; each is a window total. The suffix now appears only on measures that are ratios. +- **Cool Breeze's warning text now reads clearly on the plan viewer's properties panel** ([#4658]) - In Lite, the Darling Viewer and the Dashboard, warning-tier text in the Cool Breeze theme (the missing-index impact percentage and per-thread skew on the plan viewer's properties panel) was a shade too light to meet the accessibility contrast guideline there. The warning color is now slightly darker by default, so that text meets the guideline on the panel and on the theme's cards, plan nodes and status bar. +- **Darling: the daily digests keep their time of day** ([#4656]) - The collector-cost digest, the fleet-sweep daily rollup and the analysis singles digest were due 24 hours after their last send and checked on an hourly tick that itself slid, so each arrived up to an hour later than the day before. Each now keeps a daily slot, and a late tick no longer moves the next day's. +- **Darling Viewer: the Plan Viewer opens when the viewer has no configuration or can't reach the store** ([#4649]) - The failure message now covers only the content area. Open Plan Viewer, View Log, Open Log Folder and About stay usable; the entries that need the store are disabled. +- **Custom Views AVG panels read from a rollup match raw to the last digit** ([#4647]) - The rollup route divided in floating point after rounding the sum, so for very large sums it could differ from raw's average by one unit in the last place. +- **Darling Viewer: restarting as administrator during an upgrade keeps its configuration** ([#4646]) - The restarted viewer read its own `--upgrade-takeover` flag as the path to darling.json and showed "darling.json not found"; it now reads its arguments the way `--test` does and passes the first launch's arguments on. +- **Collectors run at their configured interval on large fleets, in Darling and Lite** ([#4644]) - Each collector's next run was scheduled from when its pass started (Darling) or waited a full minute after the pass's work (Lite), so late starts accumulated: busy Darling stores ran 1-minute collectors about every 71 seconds and lost about 15% of per-minute samples. Both now keep a fixed grid and skip, rather than replay, a slot missed during a stall. Stores that were drifting collect correspondingly more data per day: up to about 20% more for the per-minute tables. +- **The plan viewer's minimap now draws the first time you open it** ([#4643]) - Opening the minimap on a freshly loaded plan tab used to show only the "Minimap" header with nothing below it, until you dragged its resize grip or closed and reopened it. It now draws the plan overview immediately. +- **Double-clicking a minimap node now centers it correctly** ([#4643]) - Double-clicking a node in the minimap to zoom to it used to leave the node off to the side once the properties panel opened, by as much as 200 pixels. It now centers in the space that's actually left after the properties panel opens. +- **Jumping from a plan warning to its operator now centers the operator** ([#4643]) - Clicking a warning header's link to jump to the operator it came from used to leave that operator off to one side, because the properties panel opening alongside it narrowed the view after the centering had already happened. It now centers in the space that's actually left once the panel opens. +- **Copy and paste no longer freeze the window for up to 9 seconds when another app holds the clipboard** ([#4634]) - In Lite, the Darling Viewer and the Dashboard, a copy or paste that found the clipboard busy tried 8 times, and each try could take about a second, so the window stopped responding for up to about 9 seconds. It now tries twice and gives up after about 2 seconds. +- **get_store_metrics no longer reports retired store objects as a failing sweep** ([#4633]) - The inventory counted every row left from an older sweep as a sweep that "is not completing", and pointed at a Warning line that does not exist. On an upgraded store those rows belong to objects the product retired, such as the superseded baselines and the frozen rollups' jobs. It now checks whether each object still exists. Retired objects appear in their own `dropped_objects` list with the time of their last row. Only an object that still exists is reported as a sweep gap, with the service-log lines that say why. The compression line also explains why `collection_health_hourly` is uncompressed instead of listing it as DISABLED. +- **The Runtime Summary no longer says a query used "100%" of a memory grant it never got** ([#4632]) - When a query's memory grant was zero, the plan viewer now says "No memory grant" instead of a fabricated "0 KB granted, 0 KB used (100%)". +- **Plan viewer orange text now reads clearly in the Light theme** ([#4632]) - In Lite and the Darling Viewer, the warning badge, missing-index impact percentage, per-thread skew text, and the node cost/elapsed/CPU/row-estimate text used a fixed orange that was too pale to read on a white background; they now use theme-aware colours that stay readable in both Light and Dark. +- **Collector status text in the status bar now reads clearly in the Light theme** ([#4632]) - When collectors were erroring or capture was down, the status bar text used a fixed orange that was too pale to read on a white background; it now uses the same theme-aware colour the plan viewer uses, readable in both Light and Dark. +- **The repro script declares each parameter once** ([#4626]) - When a batch's plan lists the same parameter, or an auto-parameterized `@0`/`@1`, for more than one statement, the repro script declared it once per statement and failed to run. Each name is now declared once; a name the statements give different types is left out with a warning, and a warning says when the statements' compiled values differ. Matches PerformanceStudio. +- **Collection health no longer calls an idle activity-only collector regressed** ([#4625]) - `waiting_tasks`, `query_snapshots`, `running_jobs`, `job_history`, `procedure_stats`, `pg_lock_stats` and `pg_session_states` store rows only while the target is busy, so three empty runs in a row is normal for them, and working collectors showed WARNING. They are now flagged as stopped only after 72 hours without a row. Errors, denials and staleness are still flagged at once. Every row flagged as regressed now carries the sentence that explains it, including the default `get_collection_health` rows, and the web Collection Health table has a Regression column. +- **The daily purge of the per-interval Query Store tables no longer scans each table in full** ([#4615]) - The two tables get an index on the column the purge and the Query Store read gate filter on, so a purge with nothing to delete no longer costs minutes on a large store (measured: a cold-cache no-op purge dropped from ~4.1s to ~0.6ms on a 4M-row table). - **A rare deadlock between a chunk drop and the alert pass's database-state check no longer fails that pass** ([#4612]) - The alert pass retries its database-state statements once when PostgreSQL picks them as a deadlock victim, the way the daily purge already retries its chunk drop. - **Copying from Lite or the Darling Viewer no longer crashes when another program holds the clipboard** ([#4600]) - Every remaining unguarded clipboard write (recommendation copy/AI-prompt buttons, MCP command/URL copy, grid/row/cell copy, chart image copy, alert detail T-SQL copy, sign-in device-code copy) now retries briefly instead of throwing when another process is holding the clipboard. - **The plan viewer no longer crashes when another program is holding the clipboard, and "Copy Query Text" now hands back the full query on a truncated single-statement plan** ([#4593]) - Copying anything from the plan viewer, blocking chain, deadlock graph, or a results grid used to crash the app if the Windows clipboard was briefly locked by another program; copies now fail quietly instead. Separately, "Copy Query Text" on a single-statement plan whose text hit SQL Server's 4,000-character cap now returns the full query the plan was captured from, instead of a copy cut off mid-word — for the Query Store and Active Queries plan sources that pass the original query text. @@ -351,7 +493,7 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. - **Web viewer timeouts now show a message and get logged** ([#4281]) - A read that ran past the store's timeout used to show a blank error. The service log had nothing, so a timeout looked the same as a bug. Now the page says "The store took too long to answer this read. Try again in a moment." Other failures show a short generic message. For both, the service log records the route, how long the read ran and what kind of failure it was. - **A major store upgrade keeps your ALTER SYSTEM settings** ([#4280]) - A major PostgreSQL upgrade of the managed store used to drop every `ALTER SYSTEM` setting. A raised `shared_buffers`, for example, went back to the default with no warning. The upgrade now checks each setting against the new version, test-starts the new store with the settings it accepts, and keeps them if that start works. A setting the new version rejects is left out with a warning that names it. If the test start fails, the store starts on its defaults instead, and a warning names each setting it dropped. The logs name settings but never show their values. The original file is kept as `postgresql.auto.conf.pre-upgrade` next to the data directory, so a dropped setting can be re-applied by hand. - **Lite says when a query window was cut short** ([#4279]) - Three MCP tools read raw tables with no rollup fallback: `get_top_queries_by_cpu`, `get_top_procedures_by_cpu` and `get_query_store_top`. Lite keeps those tables 30 days by default, and a user can lower that per collector. So a long window was read from much less history, with nothing to say so. Each tool now returns `effective_start`, `effective_hours_back` and `window_truncated`, plus a `truncation_note` when the window was cut short. The three grids show "Showing since" and the effective start above the grid in the same case, after a time-range change or a slicer drag. On Lite and Darling alike, each tool's short description now explains `window_truncated`. The false line that said Lite always reads the full window is gone. -- **Top-CPU reads disclose a short window** ([#4278]) - Two MCP tools read raw tables that the store purges at 4 days with the rollups on. They are `get_top_queries_by_cpu` and `get_top_procedures_by_cpu`. A longer window came back short, and `hours_back` came back unchanged. Both tools now return `effective_start`, `effective_hours_back` and `window_truncated`, plus a `truncation_note`, like `get_query_store_top`. The desktop viewer's Top Queries, Top Procedures and Query Store grids show "Showing since" and the effective start in the same case. The web viewer shows the same note on those panels. +- **Top-CPU reads disclose a short window** ([#4278]) - Two MCP tools read raw tables that the store purges at 4 days with the rollups on. They are `get_top_queries_by_cpu` and `get_top_procedures_by_cpu`. A longer window came back short, and `hours_back` came back unchanged. Both tools now return `effective_start`, `effective_hours_back` and `window_truncated`, plus a `truncation_note`, like `get_query_store_top`. The desktop viewer's Top Queries, Top Procedures and Query Store grids show "Showing since" and the effective start in the same case. The web viewer shows the same note on those panels. On the hourly tier, `window_truncated` and `effective_start` come from a per-server check of where the rollup starts. When the window ends past the rollup's materialization ceiling, the note says nothing after it was read, or that the end edge is not verified when the ceiling is unknown. - **`get_query_store_top` stays under the MCP response-size budget** ([#4198]): The default call returned about 48 KB, over the 32 KB limit some MCP clients enforce. `query_text` is now a 400-character preview by default. A new `full_text` argument returns the whole statement. A new `query_text_truncated` flag on each row says whether the text was cut. The web viewer is unchanged and still shows up to 2,000 characters. - **`describe_custom_view_catalog` stays under the MCP response budget** ([#4198]): The Custom Views compose catalog returned 98 KB by default, three times the budget. It now returns a compact, source-grouped list by default (name, purpose, aggregate/unit vocabulary only). A new `source` argument drills into one source's full detail. A new `full_detail` argument returns the complete catalog as before. The web Custom Views editor is unaffected: it reads the full catalog through its own endpoint. - **MCP `get_collection_health` caps its response size** ([#4198]): A default call on a busy SQL Server fleet reached 41,669 bytes, over the 32 KB limit. Healthy collectors with nothing to report now compact to a shorter row by default. Any failing, stale, stopped, erroring, denied, or regressed collector keeps every field. A new `full_detail` argument restores every field on request. The web dashboard and Custom Views are unaffected. @@ -4221,8 +4363,112 @@ Full entries: [docs/changelog/3.0.md](docs/changelog/3.0.md) [#4599]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4599 [#4600]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4600 [#4601]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4601 +[#4602]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4602 [#4603]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4603 [#4604]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4604 [#4610]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4610 [#4612]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4612 [#4614]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4614 +[#4615]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4615 +[#4616]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4616 +[#4617]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4617 +[#4618]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4618 +[#4625]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4625 +[#4626]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4626 +[#4632]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4632 +[#4633]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4633 +[#4634]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4634 +[#4643]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4643 +[#4644]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4644 +[#4646]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4646 +[#4647]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4647 +[#4649]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4649 +[#4656]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4656 +[#4658]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4658 +[#4664]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4664 +[#4665]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4665 +[#4667]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4667 +[#4669]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4669 +[#4671]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4671 +[#4680]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4680 +[#4682]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4682 +[#4685]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4685 +[#4688]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4688 +[#4698]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4698 +[#4700]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4700 +[#4703]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4703 +[#4704]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4704 +[#4705]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4705 +[#4712]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4712 +[#4713]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4713 +[#4714]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4714 +[#4717]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4717 +[#4718]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4718 +[#4725]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4725 +[#4738]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4738 +[#4739]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4739 +[#4740]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4740 +[#4742]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4742 +[#4753]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4753 +[#4762]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4762 +[#4763]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4763 +[#4775]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4775 +[#4777]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4777 +[#4778]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4778 +[#4779]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4779 +[#4780]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4780 +[#4781]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4781 +[#4783]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4783 +[#4784]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4784 +[#4786]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4786 +[#4787]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4787 +[#4790]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4790 +[#4791]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4791 +[#4792]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4792 +[#4794]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4794 +[#4796]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4796 +[#4797]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4797 +[#4798]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4798 +[#4799]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4799 +[#4801]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4801 +[#4802]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4802 +[#4803]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4803 +[#4804]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4804 +[#4805]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4805 +[#4806]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4806 +[#4808]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4808 +[#4810]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4810 +[#4811]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4811 +[#4813]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4813 +[#4814]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4814 +[#4816]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4816 +[#4818]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4818 +[#4820]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4820 +[#4826]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4826 +[#4827]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4827 +[#4828]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4828 +[#4829]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4829 +[#4830]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4830 +[#4831]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4831 +[#4832]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4832 +[#4833]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4833 +[#4835]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4835 +[#4837]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4837 +[#4838]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4838 +[#4839]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4839 +[#4840]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4840 +[#4841]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4841 +[#4846]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4846 +[#4847]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4847 +[#4848]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4848 +[#4849]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4849 +[#4850]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4850 +[#4851]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4851 +[#4852]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4852 +[#4853]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4853 +[#4854]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4854 +[#4855]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4855 +[#4858]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4858 +[#4859]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4859 +[#4861]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4861 +[#4867]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4867 From aea1d338a10eb961835fb7ec214f55eb26938a7a Mon Sep 17 00:00:00 2001 From: Erik Darling <2136037+erikdarlingdata@users.noreply.github.com> Date: Wed, 30 Sep 2026 14:40:52 -0400 Subject: [PATCH 2/2] CHANGELOG: splice #4863 (managed store random_page_cost = 1.1) into [Unreleased] Changed --- CHANGELOG.md | 2 ++ 1 file changed, 2 insertions(+) diff --git a/CHANGELOG.md b/CHANGELOG.md index d43bbac7e..4cdd461b9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -122,6 +122,7 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. ### Changed +- **The managed store sets `random_page_cost = 1.1`** ([#4863]) - the managed PostgreSQL store now renders `random_page_cost = 1.1`, a value suited to SSD storage, in place of PostgreSQL's default of 4, which prices a random page read at four sequential ones. The planner now picks index and bitmap scans that the default priced out. Measured on a production store over an hour at each value, back to back, total statement time was 2.3% lower at 1.1, with some statements faster and a few slower. A store upgrading from 3.8 moves its settings into the managed file at the upgrade start and gets the line there, and the value is in force from the service's next start after that. A `random_page_cost` set with `ALTER SYSTEM` still wins, and one already set in `postgresql.conf` when the store upgrades is carried into the managed file unchanged. - **Per-server Query Store reads fetch fewer rows** ([#4861]) - This covers the Queries grid, MCP's Query Store top read and the Trends Query Store chart. Each filtered one server's rows of the per-interval table by collection time only, which no index held. So it fetched every stored row of that server. Each read now also bounds first_execution_time, which the table's unique key holds, so older rows are dropped before they are fetched. Results do not change. On a large store, a day-long MCP read went from 3.1 s to 0.6 s. - **SQL Server blocking and deadlock baselines count the hours collection covered** ([#4853]) - each hour-of-week's mean now divides by the days collection ran in that hour. It used to divide by the days that had events. A covered hour with no blocking or deadlocks is a measured zero, so a spike into it says the baseline measured zero instead of "first occurrence". This needs at least one captured event in the 30 days, so a server that cannot capture the events keeps no baseline. Hours with occasional events get lower means. More hours are trusted, and a trusted hour must also pass the rate multiple, so those hours alert less often. An hour whose events cluster on a few days can alert at a lower event rate than before. These baselines are now computed once per UTC day, like the CPU and I/O latency baselines. Events earlier on the current day are not part of the baseline. - **The plan viewer's Runtime Summary lists CE model above Optimization, with Early abort under it** ([#4849]) - this matches PerformanceStudio. The Early abort label is indented when an Optimization row is shown. @@ -4471,4 +4472,5 @@ Full entries: [docs/changelog/3.0.md](docs/changelog/3.0.md) [#4858]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4858 [#4859]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4859 [#4861]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4861 +[#4863]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4863 [#4867]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4867