diff --git a/CHANGELOG.md b/CHANGELOG.md index e34dcd48e..16a02c6ab 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -23,6 +23,24 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. ### Added +- **Self-monitoring alerts (Collection Stopped, Capture Down, Collector Cost Regression) open a notebook with the server's collection health and the collector's log and cost** ([#4437]) +- **Analysis-finding alerts open a notebook with that finding's summary and evidence** ([#4436]) +- **Custom-rule alerts open a notebook that charts the rule's own measure with its threshold or band** ([#4433]) +- **Poison Wait alerts open an authored notebook: top waits on both engines, waiting tasks, resource-semaphore/memory-grant detail when the wait is RESOURCE_SEMAPHORE, and a trend bound to the firing wait type.** ([#4424]) +- **Server Unreachable/Restored and the Agent-job alerts (Failed Agent Job, Long-Running Job, Agent Not Running) now open an authored notebook instead of the mechanical section-list fallback, and the collection-log web read can filter by status.** ([#4422]) +- **High CPU alerts open an authored notebook: SQL Server and PostgreSQL CPU timelines, top queries/procedures by CPU, and scheduler pressure, instead of the mechanical read list.** ([#4420]) +- **PostgreSQL Wraparound Risk, Vacuum Horizon Blocked and Replication Slot Retention alerts open an authored notebook: their own drill-down reads and a 24h trend panel, instead of the mechanical fallback section list.** ([#4419]) +- **Long-Running Query and Forced Plan Failing alerts open an authored notebook: active/completed queries and a completion-duration timeline for Long-Running Query, plan corrections and a corrections timeline for Forced Plan Failing.** ([#4418]) +- **Authored Blocking and Deadlocks alert-notebook templates** ([#4395]) - `/api/alert-notebook` now returns a forensic, multi-cell layout for Blocking Detected, Blocking Wait Time, and Deadlocks Detected instead of the flat mechanical read list every other alert type gets: a blocked-process-reports or deadlocks timeline with a cross-annotation, ranked breakdowns by object/database/lock-mode, and the existing header/status/drill-down reads, all bound to the alert's own window and database. +- **Aurora per-query peak memory as a context fact** ([#4370]) - PostgreSQL-target top-statement findings now state a statement's peak executor memory (Aurora's `aurora_stat_statements()` only) as context when available, alongside its share of window time; it never changes a finding's severity, rank, or whether it fires, and reads NULL rather than 0 off Aurora. +- **Alert-notebook read-only render on `#/triage`** ([#4368]) - the alert deep link can now render the bound notebook document returned by `/api/alert-notebook` (server, status, and forensic read cells for the firing), read-only with no fleet scope picker; falls back to the existing triage assembly when the endpoint isn't present. +- **Notebook `read` cell type** ([#4367]) - Notebooks can now include a `read` cell (`{type:"read", read, params, viz, title}`), validated by the same rules as a dashboard's v1 read panel and rendered through the shared panel renderer. Those shared rules now also check a read panel's `params` on save, for dashboards as well, so a saved dashboard panel with a malformed `params` object is rejected the next time it is saved. +- **Alert-notebook binding endpoint** ([#4366]) - added `GET /api/alert-notebook`, which turns an alert link into a bound, read-only notebook definition (a dedup-matched alert, a four-arm live status, and a mechanical per-metric read cell list), the server half of the alert-to-notebook MVP (#4222). +- **A shell-level pause/resume control for the web viewer's auto-refresh** ([#4365]) - the 60s refresh loop previously only paused automatically while a browser tab was hidden. Added a button in the sidebar chrome, present on every page, that lets the operator pause and resume the tick directly; the paused state persists across reloads. Resuming runs one refresh immediately. Also added the alert-notebook triage deep-link route to the existing poll-skip guard so the periodic tick does not re-render it. +- **The store host profile now reports its cloud instance type** ([#4330]) - `get_store_host`, `--check-settings` and the store host web panel now show the EC2 instance type or Azure VM size of the machine the Darling service runs on. For a managed store, or a store on the same machine, that is the store's own host. A short, bounded probe of the cloud metadata endpoint reads it, as the metadata service reports it, and the panel says "not detected" off the cloud or when the probe finds nothing. +- **Store host visibility: an MCP read and a web panel** ([#4282]) - The managed PostgreSQL store's own host +- **`--check-settings` reports whether a managed store's sizing still matches this host** ([#4271]) - a new CLI verb prints the host and store facts (RAM, CPUs, PostgreSQL/TimescaleDB versions, store size, buffer hit ratio, uncompressed TimescaleDB chunk bytes against RAM) and a verdict for each of the eight sizing-relevant settings: matches, stale after a hardware change, operator override, or not managed. `--json` prints the same facts for scripting. Separate exit codes for a config error, an unreachable store, and a stale setting let an install or upgrade script gate on the outcome that matters to it. +- **`--validate-config` (`--test-connection`) no longer fails a healthy store over a stale entry in darling.json** ([#4271]) - once the store is reachable, this now probes its own registry of monitored servers instead of the config file's first-bootstrap seed list. A server that is only in the file (never registered, or since removed from the registry) is now a warning, not a failure. A store that cannot be reached still falls back to the file's list, and says so on the console. This closes the false failure reported on a 42-server production store whose darling.json still listed two long-removed servers. - **Darling's MCP host now serves a /core endpoint alongside /** ([#4121]) - / keeps serving every Darling MCP tool, unchanged. /core serves a fixed 74-tool subset: the four tools a fresh conversation starts from (list_servers, get_fleet_overview, analyze_server, get_tool_guide) plus every tool the analysis engine's next_tools recommendations can point at. /core sits behind the exact same Host, bearer-token and CIDR checks as /, and a tool outside the subset fails the same way an unknown tool name would. A /core session's shared instructions now lead with a short note. The note says /core serves 74 tools, and that the tool count and tool notes after it describe the full set on /. It also says a tool outside /core fails as an unknown tool, though get_tool_guide can still describe it. Along the way, fixed eight next_tools entries that recommended get_blocked_process_reports, a tool Darling has never had (Darling's name is get_blocking), so those findings always pointed at a call that failed. - **get_resource_semaphore also returns interval_seconds** ([#4103]) - On Darling and Lite, get_resource_semaphore now returns interval_seconds beside sample_interval_seconds. It carries the sample interval under the name get_latch_stats and get_spinlock_stats already use, and it is null when the interval cannot be known. sample_interval_seconds stays for existing readers. - **get_query_store_top takes module_name, so one stored procedure's Query Store history can be read without widening top** ([#4057], contributed by @kendra-little) - the read ranked the whole window before capping it, so a procedure an investigation named but whose queries sat outside the top N could not be reached at all. The filter is part of the query on both SKUs. It matches the exact, case-sensitive, schema-qualified name the collector records (the full_name get_top_procedures_by_cpu returns, Adhoc for ad-hoc statements), and it applies after the #1841 interval dedup and before the ranking and the cap. Every row now carries module_name, beside the execution_type #4060 added, and the two filters compose. A filter that matches nothing answers empty only when the same read without the filters has rows, and the message names both filters. On Darling, a module miss also carries effective_start, effective_hours_back and window_truncated as hints. Otherwise the read answers as an unfiltered one would. @@ -87,6 +105,26 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. ### Changed +- **Report alerts link to their own page: the Fleet Sweep Rollup opens the sweeps page, and the digests carry no link** ([#4421]) +- **`get_top_procedures_by_cpu` and the Top Procedures grid now route to the hourly rollup once raw ages past its retention window, instead of returning nothing** ([#4413]) - Mirrors the routing already shipped for `get_top_queries_by_cpu`; the payload discloses which tier answered and what an hourly-routed row is missing (`object_type`). +- **Sized the Linux compose store's background-worker slots and `work_mem` from the product's own +- **Cache the server-scoped watermark read across a runner's lifetime, cutting repeated `MAX()` reads against `job_history`, `default_trace_events`, `system_health_events`, and `memory_pressure_events` to one seed per (server, collector) pair instead of one per collection cycle** ([#4399]) +- **Top-N-by-CPU queries now route to the hourly rollup once raw's retention has dropped the window, instead of returning empty** ([#4396]) - `get_top_queries_by_cpu` (MCP tool, Storage reader, and Viewer) degrades to the hourly continuous aggregate for a window older than raw's floor and discloses the precision loss (`tier_used`, `precision_note`) rather than silently returning nothing. +- **The Performance Trends Query Store duration chart reads the interval table for windows of 48 hours or more** ([#4382]) - the table measured faster than raw from 48 hours (361 ms vs 423 ms at 48 h; 374 ms vs 1,045 ms at 7 days) and returns exactly raw's points; shorter windows stay on raw. The interval-table gate also reads raw for any window that still holds legacy rows without an interval start. +- **The Darling viewer's TempDB file I/O trend now buckets like the File I/O tab's own reads** ([#4353]) - a 7-day window used to return one row per collection per tempdb file (40,324 rows measured for 4 files at 1-minute cadence); it now buckets to the chart's point budget (4,036 rows measured for the same seed) and still stamps a short window's raw points unchanged. +- **Raw hypertables now re-tune their own chunk interval once a day from actual ingest** ([#4344]) - Every raw hypertable used the same fixed one-day chunk interval whatever its write rate, so a busy table's open chunk, with its indexes, could outgrow the memory meant to hold it. Darling now reads each raw hypertable's compressed ingest rate daily, alongside a RAM-derived budget, and narrows the interval one rung at a time through `set_chunk_time_interval` when the store-wide total is over budget, and widens it one rung when the total would stay under half the budget; a new rung history table records every change, alongside a per-run WAL-volume record. On by default; set `rawChunkIntervalReconcileEnabled: false` in the config file to turn it off. +- **The Queries grid and MCP's Query Store top read use a per-interval table for long windows** ([#4341]) - both used to deduplicate the whole raw Query Store slice on every read. The grid now reads a table that keeps one row per interval for windows of 12 hours or more, and the MCP read for 24 hours or more. On a 15-day seed of 3.47 million raw rows, a 7-day window took 2,754 ms against 7,065 ms for the grid, and 1,627 ms against 9,249 ms for the MCP read. Shorter windows, and any window the table's history doesn't cover yet (after an upgrade), read raw as before. +- **Lite's memory clerk and File I/O trend charts bucket long windows in DuckDB** ([#4340]) - The memory clerk, File I/O latency, File I/O throughput and TempDB File I/O trend charts, and the Overview I/O lane, read one point per collection. They now gather collections into buckets sized to the window, at most 1,500 points per series. A bucket's latency is its summed stall over its summed operations, and its throughput is its summed bytes over its summed seconds. A window short enough that every bucket holds one collection plots each point at its own collection time, as before. The memory clerk picker's name list is cached for 15 minutes. +- **Lite keeps one DuckDB connection open, so reads attach to it instead of reopening the database file** ([#4339]) - Every read used to open and close its own connection to the local database file. One connection now stays open for the app's life, which cut a 20-tick, ten-way Overview read on a 1 GB store from 871 ms to 43 ms. Because that connection keeps DuckDB's buffer pool resident, a periodic check trims it after a heavy read: on a 1.1 GB store, one trim took process private bytes from 1,115 MB to 148 MB. +- **Lite's Performance Trends charts bucket long windows in DuckDB** ([#4338]) - The query duration, procedure duration and execution count charts read one point per collection, about 10,080 points each for a 7-day window at one collection a minute. They now gather collections into buckets sized to the window, rating each bucket as its summed work over its summed seconds, and a window short enough that every bucket holds one collection plots each point at its own collection time, as before. +- **Lite's Overview CPU, wait and memory lanes bucket long windows in DuckDB** ([#4337]) - The Overview tab's CPU, wait and memory lines, and the CPU and Memory tab charts, read one point per collection: about 10,080 points each for a 7-day window at one sample a minute. They now gather collections into at most 1,500 points per line, and a window short enough that every bucket holds one sample or collection plots each point at its own collection time, as before. +- **Settings the service manages now live in one included file, not stacked append blocks** ([#4336]) - the v1-v15 `postgresql.conf` append blocks are replaced by a single included file, `darling-managed.conf`, rewritten wholesale on every start. Existing stores migrate their current effective values into it on the first start after upgrade, without changing any value. An operator's own lines are kept, moved below the include so they keep winning, and settings set with `ALTER SYSTEM` are never changed. From then on, sizing values can change on the first start after an upgrade, and each change is logged with the inputs that produced it. A failed verification restores the previous settings files and raises the store-settings alert. Operator lines placed below the include are not yet carried across a major PostgreSQL version upgrade ([#4358]). +- **Darling viewer's Performance Trends charts bucket server-side over long windows** ([#4333]) - The query duration, procedure duration and execution count charts read one point per collection on the raw tier, about 10,080 points each for a 7-day window at one collection a minute. The store now gathers collections into buckets sized to the window, rating each bucket as its summed work over its summed seconds, and the chart title names the bucket width. A window narrow enough that every bucket holds one collection plots each point at its own collection time, as before. +- **Lite's Wait Stats and Perfmon charts bucket long windows instead of shipping one point per collection** ([#4331]) - A 7-day Wait Stats or Perfmon chart carried one row per stored collection, about 10,080 rows per wait type or counter at one collection a minute (about 200,000 for 20 wait types). It now buckets each wait type or counter to at most about 1,500 points (the chart budget, not the MCP tools' 200-point budget), and a window short enough that every bucket holds one collection still plots each point at its own collection time. The wait-type and perfmon-counter pickers also cache their name list for 15 minutes per server and window length instead of re-reading it on every 1-minute auto-refresh. +- **Darling viewer's File I/O and memory clerk charts load faster over long windows** ([#4329]) - the +- **Darling viewer's CPU, Overview wait and Overview memory charts bucket server-side over long windows** ([#4327]) - The CPU chart, and the Overview tab's total-wait and buffer-pool lanes, used to return one row per ring-buffer sample or per collection with no cap: about 10,080 rows each for a 7-day window at one collection a minute. Postgres now buckets all three to a fixed row budget, matching the sizing already applied to the Wait Stats and Perfmon trend charts. A window narrow enough that every bucket holds exactly one sample or collection still renders each point at its own exact timestamp rather than a bucket boundary, so short windows look unchanged. +- **PostgreSQL-target baselines recompute once a day instead of every hour** ([#4324]) - Every PostgreSQL-target baseline arm (13 unkeyed metrics plus the 2 statement-keyed ones) was recomputing its full 30-day window every hour; one of them cost up to 2.45 seconds and 262 MB of temp space per pass. They now share the once-a-day cache lifetime the two heaviest SQL Server baselines (Cpu, IoLatency) already used. Because a daily baseline covers history up to the start of the UTC day, a PostgreSQL target added today has no baselines, and its baseline-driven anomaly checks sit out, until the next UTC midnight. +- **Three legacy query-stats rollups stop refreshing** ([#4186]) - `query_stats_hourly`, `procedure_stats_hourly`, `query_stats_db_hourly` and their daily rollups keep the history they already hold. Their interval-honest successors take over. Reads older than the successors' own history still use them. New data goes only to the successors, so the store no longer spends refresh time on the old rollups. The raw `query_stats` and `procedure_stats` purge now waits until the successors and the old rollups together cover every raw row. After a long outage, the service fills the gap at startup, up to a day of it per start. Only then does the purge resume. `--backfill-rollups` fills it in one run. Existing compression jobs move to new hours once. Dashboards, alerts and queries do not change. - **llms.txt and CITATION.cff now match the shipped product** ([#4157]) - llms.txt said 41 T-SQL collectors (it is 42), listed Azure SQL Database under Lite only (Darling supports it too), never mentioned Darling's 29 PostgreSQL collectors or its Windows/Linux and web-dashboard support, named only email and tray alerts (Teams, Slack, PagerDuty and generic webhooks also fire), and undercounted downloads (15,000+, not 4,300+). SentryOne is now named as SolarWinds SQL Sentry. CITATION.cff drops a version and release date no release step keeps current. A new test compares both collector counts in llms.txt against the collector catalog so they cannot go stale silently again. - **The MCP tool list's size limit now matches its size** ([#4141]) - #3898 moved reading guidance out of the tool list. The budget test's total limit now equals the tool list's measured size. On Darling that is 170,798 bytes, down from a limit of 329,494. On Lite it is 89,719 bytes, down from 148,216. Any later growth has to raise the limit on purpose. The window_truncated note on the time-series tools also drops its issue number, and now says the field was formerly named truncated. - **get_spinlock_stats puts a short description in the tool list** ([#4127]) - Before this change, this Darling and Lite MCP tool put 597 and 1,306 characters into the tool list. It now serves a head of 595 characters on both, under the usual 620-character cap. The get_tool_guide tool returns the rest of the guide by tool name. Nothing was deleted, and no tool behavior changed. @@ -168,6 +206,96 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. ### Fixed +- **PostgreSQL statement statistics count a re-created statement entry's work since it was re-created, instead of under-counting it; on PostgreSQL 16, where the entry's creation time isn't available, such a row is recorded as unknown** ([#4435]) +- **Wait statistics on servers where a scheduled job clears them are counted from each clear instead of being recorded as unknown; only the work between the last collection and the clear stays unknown** ([#4434]) +- **The raw purge gate judges each rollup against the rows that rollup can hold, so an hour of CPU-unknown query statistics at the floor can't hold the purge forever** ([#4432]) +- **Query statistics record a restarted cached plan's executions, duration and CPU since the restart in full, instead of under-counting them; a restart the collector cannot place in time is recorded as unknown rather than estimated** ([#4431]) +- **After an outage across an upgrade, the raw purge resumes within the hour on a running store, with no second restart** ([#4430]) - a legacy/successor seam an outage opens across an upgrade used to require a second full service start before the raw purge gate would release; the hourly retention tick now closes that seam on its own. +- **The daily retention sweep no longer drops raw query statistics over a hole no rollup holds; the gated service-triggered purge owns those tables, and purge_now reports what it held** ([#4429]) +- **Query statistics no longer record a false zero CPU when only the CPU counter's delta is unknowable; those rows keep their executions and duration** ([#4423]) - a `query_stats` row whose CPU (worker time) counter reset or was seen for the first time, while its execution-count or elapsed-time counters had a real delta in the same pass, previously wrote CPU as a false 0 over a fabricated zero interval — which then caused the row's real executions and duration to be dropped by the interval-honest filter along with the fake CPU. That row now records CPU as unknown (NULL) and keeps its real interval, executions, and duration. A CPU counter that is genuinely unchanged over a real measured interval still records 0, unchanged. +- **FinOps no longer reports a database as idle when its only recent samples were a plan's first sighting** ([#4417]) - idle-database detection now counts a recent last-execution time as activity, instead of deciding on the summed execution count alone (#4394). +- **Alert notebook: the collector-freshness check now counts only successful collection runs** ([#4416]) - it previously treated a collector's own failed runs as proof it was still fresh, used no read timeout, and always logged a zero elapsed time on failure. It also now recognizes AG and Server Unreachable/Restored recovery pairs when resolving a notebook's status. +- **`get_store_host` now honours request cancellation** ([#4412]) - an abandoned or superseded call to `get_store_host` kept its store-host gather running to the tool's own fixed deadline instead of stopping when the caller went away; the gather now also observes the request's own cancellation, and a cancelled gather is never cached. +- **Web reads honour request cancellation for 2 more trend charts and 3 analysis tools** ([#4411]) - `get_file_io_trend`, `get_perfmon_trend`, `compare_analysis`, `get_analysis_facts` and `get_analysis_findings` kept running their store reads (and, for the analysis tools, the inference engine's fact collection and scoring) to completion after a web request was abandoned. They now take a cancellation token and stop when the client disconnects. +- **Web reads for PostgreSQL index bloat, column stats, CPU utilization, log events and logging audit now honour request cancellation** ([#4409]) - an abandoned web request for one of these five reads previously ran its underlying store query to completion instead of stopping; each now threads the request's cancellation token through every store call. +- **The settings redactor no longer masks a short value whole on its first call** ([#4408]) - its patterns are warmed once at start-up, so a first-use compile cost can't count against the match time bound. +- **darling-managed.conf is rewritten only when its settings change** ([#4407]) - a change in a header display field, such as the data volume's free space crossing a GiB boundary, no longer rewrites the file or re-verifies it. +- **Stop re-holding the raw retention purge for a healthy store whose successor rollup hasn't run its first refresh yet** ([#4406]) - An upgrade with no outage was reporting the raw purge's safety gate as unsafe while a freshly-added rollup's history was still empty, even though no row was actually at risk. The gate's fallback probe now stops one refresh-window short of the current time, matching how far the rollup's own first refresh is expected to reach. +- **Carry an operator's own `postgresql.conf` lines below the `darling-managed.conf` include across a major PostgreSQL upgrade** ([#4405]) - previously only `postgresql.auto.conf` (`ALTER SYSTEM`) settings survived a major upgrade; a hand-edited line below the include stayed behind in the retained old data directory. Each carried line is now probed against the new binaries and skipped (never fatal) if rejected. +- **Web reads honour request cancellation (10 more tools)** ([#4402]) - `get_store_query_stats`, `get_store_metrics`, `get_store_log`, `get_collector_stall_probes`, `get_plan_xml`, `get_oversized_plan_backlog`, `get_fleet_overview`, `get_sweep_reports`, `get_collector_cost` and `get_pg_blocking` kept running their store query to completion after a `/api/read` client disconnected or timed out, because their dispatch entries and tool methods never threaded the request's `CancellationToken` through to the query. Each now takes the token and passes it to every store call, so an abandoned request stops promptly instead of finishing unread. +- **Fixed a raw-purge gate that could hold history forever, and a repair +- **Web reads honour request cancellation (11 more SQL Server tools)** ([#4393]) - `get_session_stats`, `get_waiting_tasks`, `get_cpu_scheduler_pressure`, `get_plan_cache_bloat`, `get_latch_stats`, `get_spinlock_stats`, `get_memory_trend`, `get_default_trace_events`, `get_running_jobs`, `get_pvs_stats`, and `get_ag_health` now take a `CancellationToken` and pass it down to every store call, so an abandoned `/api/read` request stops its query instead of running to completion. Removed all from `CancellationAllowlist`. +- **The PostgreSQL settings redactor has a bounded matching time** ([#4392]) - Each redaction pattern runs under a time bound. A value that reaches it is masked whole, and the warning names only the setting. +- **An armed raw-retention purge no longer runs when PostgreSQL starts, before the service can hold it** ([#4391]) - After a long enough outage, TimescaleDB's own scheduler could run the `query_stats`, `procedure_stats` or `query_store_stats` retention job the moment PostgreSQL came up, dropping raw history that no rollup had materialized yet. These three jobs are now never scheduled on TimescaleDB's runner. The service runs the purge itself, from its hourly evaluation only, and only after it measures coverage itself, confirms the hole repair finished cleanly since the current PostgreSQL start, and finds no hole in the range it would drop. A new **Raw Purge Over Horizon** alert says why a purge didn't run, or that the purge trigger has stopped. The first start after the upgrade still runs the old armed job once. +- **PostgreSQL web reads now honour request cancellation** ([#4390]) - `get_pg_xmin_horizon`, `get_pg_wraparound_risk`, `get_pg_session_states`, `get_pg_replication_stats`, `get_pg_predicate_stats`, `get_pg_kernel_stats`, `get_pg_io_stats`, `get_pg_index_usage`, `get_pg_database_stats` and `get_pg_autovacuum_health` ran every store call to completion even after the web caller disconnected. Each now takes a `CancellationToken` and passes it to every store call, so an abandoned request stops its query instead of running to completion. +- **PostgreSQL-target MCP reads now cancel on client disconnect** ([#4388]) - nine PostgreSQL-target read tools (`get_pg_deadlocks`, `get_pg_deadlock_detail`, `get_pg_wait_stats`, `get_pg_wait_sampling`, `get_pg_replication_slots`, `get_pg_top_queries`, `get_pg_table_bloat`, `get_pg_plans`, `get_pg_plan_capture_readiness`) ran to completion even after the requesting client disconnected. Each now takes a `CancellationToken` threaded to every store call and is dropped from the web-read cancellation allowlist. +- **Web reads now honour request cancellation for 15 more MCP tools** ([#4387]) - `get_cpu_utilization`, `get_wait_stats`, `get_wait_types`, `get_wait_trend`, `get_memory_stats`, `get_memory_clerks`, `get_file_io_stats`, `get_tempdb_trend`, `get_perfmon_stats`, `get_query_store_top`, `list_servers`, `get_collection_health`, `get_collection_log`, `get_current_waits_trend` and `get_blocking_stats` used to run their store query to completion even after the web viewer's request was aborted; each now takes a `CancellationToken` threaded to every store call, so an abandoned request stops promptly. +- **Web reads honour request cancellation for Alerts, Memory Grants and Health MCP tools** ([#4386]) - `get_alert_history`, `get_alert_settings`, `get_mute_rules`, `get_notification_routes`, `get_resource_semaphore`, `get_memory_grants`, `get_memory_pressure_events`, `get_server_summary`, `get_daily_summary` and `get_daily_summary_range` previously ran their store query to completion even after an abandoned web request; they now take a `CancellationToken` and pass it through every store call. +- **Threaded cancellation through 10 PostgreSQL-target web reads** ([#4373]) - get_pg_buffer_usage, get_pg_extensions, get_pg_lock_stats, get_pg_server_config, get_pg_server_config_changes, get_pg_write_stats, get_pg_database_trend, get_pg_io_trend, get_pg_query_duration_trend and get_pg_wait_trend ran their store queries to completion after a web request was abandoned; each tool now takes a CancellationToken and passes it to every store call, and is removed from the /api/read cancellation allowlist. +- **Web reads for object-stats and config-history now honor request cancellation** ([#4372]) - get_database_sizes, get_index_usage, get_object_locking, get_table_index_sizes, get_database_config_changes, get_database_scoped_config, get_server_config_changes and get_trace_flag_changes ran their store query to completion even after a web caller abandoned the request; all eight now take a CancellationToken threaded through every store call and are removed from the cancellation allowlist. +- **Deprecated Dashboard's XE ring-buffer collectors no longer reshred unchanged data** ([#4371]) - `install/22_collect_blocked_processes.sql` and `install/24_collect_deadlock_xml.sql` converted the whole Extended Events ring buffer to XML and shredded it on every scheduled run, even when nothing new had arrived since the last cycle. Both now read the ring buffer target's own delivered-event counter first and skip the conversion when it hasn't changed, matching the fix #4212 shipped for Darling and Lite. Applies to existing installs still running the deprecated Dashboard's SQL Agent collectors; the rows each collector stores are unchanged. A no-change collector run (median of 5, on a filled ring buffer) went from 150 ms to 35 ms for blocked-process and from 56 ms to 28 ms for deadlock. +- **Bucketed the memory-grant and TempDB usage viewer trend charts** ([#4364]) - The memory-grant overlay, Memory Grants chart, and TempDB usage chart (Darling and Lite) each returned one row per collection over a 7-day window. On a 7-day seed at each collector's cadence, the memory-grant overlay went from 10,080 to 1,008 rows, the Memory Grants chart from 20,160 to 2,016, and TempDB usage from 10,080 to 1,008. All three now bucket to the chart's point budget the same way the CPU, wait, and file-I/O trend charts already do, averaging gauges and summing true deltas per bucket, with no change to what the series measures. +- **Bucketed the blocking-trend charts' lock-wait, waiting-task, and blocked-session reads** ([#4362]) - the Blocking Trends and Current Waits charts read one row per collection at wide windows, unlike the other trend charts #4353 and #4340 already bucketed. On a 7-day seed at the 1-minute cadence, lock waits went from 20,160 to 2,016 rows, waiting tasks from 30,240 to 3,024, and blocked sessions from 20,160 to 2,016, on both products. Blocked sessions is now averaged per collection within each bucket, so its value no longer grows with the bucket width. +- **Bucketed the CPU scheduler, session stats, and plan cache trend reads** ([#4361]) - These three reads returned one row per collection with no cap, unlike the file I/O, memory, wait stats, and other trend reads #4234 already fixed. On a 7-day seed at each collector's cadence, CPU scheduler pressure went from 10,080 to 1,008 rows, and Sessions and Plan cache each from 2,016 to 1,008. All three now bucket to the same width ladder the other trend charts use, averaging each gauge per bucket and (for Session Stats) carrying forward the newest collection's application/host attribution rather than blending it. +- **Blocking and deadlock web reads now cancel with the request** ([#4203]) - the seven blocking/deadlock tools behind the web viewer's Blocking page (blocked process XML, blocking, blocking trend, deadlock detail, deadlock trend, deadlocks, lock wait trend) kept running their store queries after a browser navigated away or a request was abandoned. Each now takes a cancellation token and passes it through every store call, so an abandoned request actually stops. +- **Health Parser web reads now cancel with the browser request** ([#4359]) - the nine `get_health_parser_*` `/api/read` handlers ran their Postgres queries to completion even after the browser tab closed or the request was aborted; they now thread `RequestAborted` through every store call so an abandoned read stops promptly. +- **Four Performance-Trends web reads now stop their store query when the request is abandoned** ([#4357]) - `get_query_trend`, `get_query_duration_trend`, `get_procedure_duration_trend` and `get_query_store_duration_trend` used to keep running their store query after a web caller gave up or navigated away; they now observe the request's cancellation all the way down to Npgsql. +- **A slow first-run database create no longer fails Darling's Postgres bootstrap** ([#4352]) - The managed store's first-run `CREATE DATABASE` was retried up to 6 times inside the same 10 second window used for the post-start connection probe. Because a timed-out `CREATE DATABASE` is cancelled and rolled back, every retry re-copied the template database from scratch under the same 10 second budget, so a copy that took longer than 10 seconds failed outright no matter how many retries ran. The create now runs once, on a 60 second budget, outside that retry. +- **Passwords and other secrets in stored PostgreSQL settings are now redacted** ([#4351]) - The config collector stored `pg_settings` and per-database/per-role override values verbatim, including secrets a `pg_monitor`-only monitoring role can read but should not retain in plain text, such as a standby's `primary_conninfo` replication password or a backup tool's key/token in `archive_command`/`restore_command`; new collections now mask the secret, and existing stored rows are scrubbed in place, in batches, and the scrub completes at any settings cadence; a target that fails is retried on the next start. Non-secret password-policy settings (`rds.accepted_password_auth_method`, `rds.restrict_password_commands`, `passwordcheck.min_password_length`) are no longer masked. "Keep the rest of the value" is not true for every case: `ssl_passphrase_command` and a dotted extension setting naming a secret are masked whole. Rotate any password or key that was ever stored this way: the scrub changes the store's current rows only, not copies made before it ran. Upgrade every service that writes to a store - a service still on an old build keeps storing plaintext after the scrub's marker is set, and nothing re-runs it. +- **Eight Queries-tab reads in the web viewer stop their store query when you leave the page** ([#4350]) - get_active_queries, get_plan_corrections, get_query_heatmap, get_query_store_clutter, get_top_queries_by_cpu, get_top_procedures_by_cpu, get_query_store_regressions and get_long_query_completions used to run their store query to the end after the browser gave up on the request. They now stop when the request is abandoned. The tab's duration trends and Query Store top list still run to the end; later PRs convert them. +- **Configuration-tab reads in the web viewer stop their store query when you leave the page** ([#4347]) - Leaving the Configuration tab, or an audit_config read, before it finished left its store query running to the end for a caller who was gone. Those six reads now cancel when the browser abandons the request. The viewer's other reads still run to the end; follow-up PRs convert them. +- **PostgreSQL `pending_restart` can now be trusted after a reload, on Windows targets** ([#4345]) - `pg_settings.pending_restart` is backend-local, so on Windows, where the collector's reconnect-every-cycle model always sees a brand new backend, a setting that had actually been changed and reloaded but still needed a restart could read `false` forever. Unix targets never had this gap: there, the postmaster itself applies the reload and every backend it forks inherits the flag, so `pg_settings` alone was already correct. The collector now also checks `pg_file_settings` on Windows targets, which is read from the file rather than a connection's own state, when the monitoring role can read it (two grants: `SELECT` on the view and `EXECUTE` on `pg_show_all_file_settings()`). `get_pg_server_config` and `get_pg_logging_audit` note when a Windows target's role cannot read it, and say what the two grants expose, so an operator can weigh the trade before granting them; a non-Windows target is untouched and never asked to widen its role for this. +- **Heal the Postgres v8 hardware-sizing block on stores resized before the #4225 fix shipped** ([#4342]) - A managed store whose hardware had not changed since its last hardware-sizing rewrite never re-triggered the heal, so a store that already carried duplicate sizing blocks, or a block written before `work_mem` rejoined the formula, stayed on stale settings even after upgrading past the earlier fix. The heal now also runs when the conf carries more than one sizing block, or when the newest block's content no longer matches what the current build would write. +- **Fleet Sweeps reports an error instead of "no sweeps" when its store read fails** ([#4328]) - A failed read for the sweep timeline, a sweep's verdicts, its would-have-paged ledger, or the watch-item worklist used to log the failure and silently answer an empty list, so the web page showed its normal "No sweeps in this span" empty state and the `get_sweep_reports` MCP tool answered an empty or silently incomplete document instead of reporting that the read failed. Both surfaces now answer the failure: the web page shows its error strip, and the tool answers its usual error status. +- **Lite Overview baseline bands used the wrong hour under a custom time range on a server not on UTC** +- **Dashboard query comparison grids and Server Trends baseline bands now use the server's local time** ([#4321]) - Under a preset time range (not a custom one), the Query Stats, Proc Stats and Query Store "Compare to" grids, and the Server Trends baseline bands, built their current time window from the Dashboard machine's clock in UTC instead of the monitored SQL Server's own local time. On a server not running in UTC, this shifted the comparison grids' "current" window and picked the wrong hour-of-day/day-of-week baseline bucket, by the server's UTC offset. Fixed to use the server's local time, matching the Server Trends ghost-line fix in [#4317]. +- **`get_collection_health`'s default reply now fits the MCP response budget** ([#4319]) - A collector row that +- **Dashboard Overview ghost line reads the right hours outside UTC** ([#4317]) - Under a preset time range, the +- **The Query Store activity slicer reads less history on each refresh** ([#4311]) - The slicer read one extra day of Query Store history below its window, as a margin for clock skew. The margin is now one hour. For the last 24 hours, it now reads 25 hours of history instead of 48. +- **Lite's Overview ghost line reads the right hours outside UTC** ([#4309]) - Under a preset time range, +- **Availability Groups tab no longer freezes the UI every 30 seconds** ([#4308]) - The Darling viewer's and Lite's Availability Groups tabs rebuilt every card and every per-database grid from scratch on every refresh. In the Darling viewer, that blocked the UI thread for over 2.5 seconds on a 42-AG fleet, even when nothing had changed. Refreshes now update existing cards in place, skip entirely when the data has not changed, and only render the cards that are actually on screen. +- **The Daily Summary calendar and `get_daily_summary_range` cache closed days for an hour** ([#4307]) - The WPF Daily Summary calendar (every 1-minute refresh) and `get_daily_summary_range` (the web Overview tab and MCP clients) recomputed every day in the displayed range on every poll, even though days before today do not change. A refresh now recomputes only today (and briefly yesterday, for a late collector run) and reuses the rest from a one-hour cache. A row that lands late for a closed day -- an outage catch-up, a backfill -- can take up to an hour to show, where it showed immediately before; an explicit `as_of` timestamp still always reads live. The cache holds at most 1,024 ranges, and drops entries older than an hour the next time it computes a range. +- **Query Store backfill's candidate check no longer reads old compressed chunks, and no longer drops a hole on a database that goes quiet** ([#4306]) - The backfill's candidate-database scan read across all of raw retention on every check, which decompressed chunks outside the backfill window: measured at 206 ms and over 21,000 buffers per check on a heavy server. It is now bounded at the backfill's own horizon (1.7 ms and 792 buffers in the design measurement). The per-database floor lookup now stops at the first row at or before that horizon instead of taking the minimum over every row. A database that stopped sending Query Store data while a repair "hole" was still open for it used to age out of the backfill scan and never get serviced. It now stays in scope until the hole is filled or expires. Both Lite and the Darling service get the change. +- **The Darling viewer's Wait Stats and Perfmon trend charts read time buckets, not every collection** ([#4304]) - Both charts read every collection in the window on each refresh: for 20 wait types over 7 days at a 1-minute cadence, that was about 170,000 rows. They now read buckets sized to the chart's point budget, which cut the rows returned about tenfold for a 7-day window in the measurement. A bucket's rate is its summed wait (or counter delta) over its summed collection interval. When the budget covers every collection, the chart still gets each raw point at its own timestamp. The wait-type and counter pickers ran a full-window DISTINCT on every 1-minute refresh; they now reuse the list for up to 15 minutes for the same server and window. +- **The Darling viewer's Active Queries grid and wait drill-down no longer load every snapshot's plan XML up front** ([#4303]) - +- **Lite's Queries-tab comparisons use the grid's time window** ([#4302]) - On a server not on UTC, the "Compare to" baseline on the Queries tab read a window shifted by the server's offset. That covers the Top Queries, Top Procedures and Query Store grids. It happened under a custom date range or after a slicer drag. Comparisons now read the same UTC window as the grid beside them. One behavior changes on purpose. After a slicer drag, the baseline is the dragged window shifted back by a day or a week, not the whole preset range. +- **Active Queries and wait drill-down no longer load every plan's XML just to show the grid** ([#4297]) - +- **The query heatmap loads much faster on large windows** ([#4295]) - The heatmap read resolved the query text of every row in the window, then kept one row per cell. Now it resolves the preview text only for the row that each cell shows. The cells, counts, top queries and preview text do not change. On a seeded window of 180,000 rows over 24 hours, the read went from 2,035 ms to 242 ms. +- **Less write-ahead log from the query, procedure and Query Store history tables** ([#4294]) - Five indexes on these tables wrote close to a quarter of the store's write-ahead log. They were read only a few times in several days. They are now dropped. Each feature that read them was timed without its index and stays well within its time limit. +- **Web viewer no longer shows raw error text for failed requests** ([#4293]) - Several pages and API routes answered a failed request with the underlying exception's own text, which could include a PostgreSQL SQLSTATE and message or a store's host and port: saving or deleting a Custom View, an alert rule or a mute rule (including turning a mute rule on or off), running a Custom View panel, the two fleet-sweep reads, a server lookup that could not read the server registry, and the triage page's notes and per-section cards. These now answer a fixed, generic message; the real text still goes to the service log so an operator can see what failed. A statement timeout now answers with a "took too long" message and the right status code in each of these places, matching what other reads already did. The one exception is a Custom View panel's own query error (a syntax error, a missing column, a bad value, or its own statement timeout), which still shows the database's message so the author can fix the panel. +- **Empty legacy baseline aggregates now drop instead of staying forever** ([#4292]) - On a store where nothing ever fed the old `perfmon_baseline` or `wait_stats_baseline` aggregate, both it and its replacement stayed empty. The empty replacement has no oldest bucket, so the old aggregate's cleanup rule never passed. Now an old aggregate that holds no rows drops on the next start, with its refresh, retention and compression jobs. +- **CPU and I/O-latency baselines recompute once a day, not every hour** ([#4291]) - Darling and Lite recomputed these two anomaly baselines every hour. Each time, they scanned 30 days of raw rows. The I/O-latency scan alone spilled about 50 MB to temp. They now recompute once per UTC day, and their 30-day window ends at midnight UTC. A 30-day baseline barely moves within a day. A failed compute still retries within the hour. +- **Fewer needless writes to the Darling store's plan and text tables** ([#4288]) - A plan or text seen again within the hour still cost a row lock and a WAL record. PostgreSQL locks the row before it checks the freshness rule. On production stores, that was 12 to 23 percent of all WAL. The upsert now skips a fresh row before it takes any lock. Separately, the touch that keeps Query Store plan and text rows from expiring re-stamped each row every 6 hours. It now waits 12 hours, which halves those writes. Rows are kept just as long as before. +- **Managed store: less WAL, from compressed full-page images** ([#4287]) - Most of a managed store's WAL was full-page images, written without compression. Every managed store now sets `wal_compression = lz4` on its next service start. In a test burst, that cut WAL by about a quarter. A longer checkpoint interval cuts WAL further, but it can slow checkpoint syncs past the store's own alert bar. So it waits for a production measurement. Bring-your-own PostgreSQL stores do not change. +- **Web failure handling: log batches, bad requests and error text** ([#4286]) - A route cut inside a character made Darling's service log drop its whole 5-second batch. The viewer's, Lite's and the Dashboard's log files had the same bug for any text with a broken character. A malformed or oversized request, or a client that dropped the connection, cost a 500 and an Error line. It now gets its own status and no Error line. A line break in exception text started what looked like a second entry in the service log. Each line is now cleaned before it is written. The read dispatcher and the fleet sweep reads answered some failures with raw exception text, such as a connection error that names the store's host. They now answer with a fixed message and log the detail. Darling's web and MCP hosts now always run as Production, so a stray Development setting on the machine cannot turn on the developer error page. +- **Time-based reads no longer shift on a store outside UTC** ([#4285]) - Every timestamp in the Darling store holds UTC, but a bring-your-own store keeps the session time zone its owner set. Reads that compare those timestamps with the current time shifted by the zone's offset. Their window was wider or narrower than asked, with no error. Every store connection that the service, the CLI and the desktop viewer open now sets its session time zone to UTC. Managed stores already set UTC in their `postgresql.conf`. The service and the viewer also keep a connection-string setting that has an empty value, such as `Password=''`, when they add their own settings. Before, they dropped it, and Npgsql then used the `PGPASSWORD` environment variable or a password file if either one was there. +- **Web viewer timeouts now show a message and get logged** ([#4281]) - A read that ran past the store's timeout used to show a blank error. The service log had nothing, so a timeout looked the same as a bug. Now the page says "The store took too long to answer this read. Try again in a moment." Other failures show a short generic message. For both, the service log records the route, how long the read ran and what kind of failure it was. +- **A major store upgrade keeps your ALTER SYSTEM settings** ([#4280]) - A major PostgreSQL upgrade of the managed store used to drop every `ALTER SYSTEM` setting. A raised `shared_buffers`, for example, went back to the default with no warning. The upgrade now checks each setting against the new version, test-starts the new store with the settings it accepts, and keeps them if that start works. A setting the new version rejects is left out with a warning that names it. If the test start fails, the store starts on its defaults instead, and a warning names each setting it dropped. The logs name settings but never show their values. The original file is kept as `postgresql.auto.conf.pre-upgrade` next to the data directory, so a dropped setting can be re-applied by hand. +- **Lite says when a query window was cut short** ([#4279]) - Three MCP tools read raw tables with no rollup fallback: `get_top_queries_by_cpu`, `get_top_procedures_by_cpu` and `get_query_store_top`. Lite keeps those tables 30 days by default, and a user can lower that per collector. So a long window was read from much less history, with nothing to say so. Each tool now returns `effective_start`, `effective_hours_back` and `window_truncated`, plus a `truncation_note` when the window was cut short. The three grids show "Showing since" and the effective start above the grid in the same case, after a time-range change or a slicer drag. On Lite and Darling alike, each tool's short description now explains `window_truncated`. The false line that said Lite always reads the full window is gone. +- **Top-CPU reads disclose a short window** ([#4278]) - Two MCP tools read raw tables that the store purges at 4 days with the rollups on. They are `get_top_queries_by_cpu` and `get_top_procedures_by_cpu`. A longer window came back short, and `hours_back` came back unchanged. Both tools now return `effective_start`, `effective_hours_back` and `window_truncated`, plus a `truncation_note`, like `get_query_store_top`. The desktop viewer's Top Queries, Top Procedures and Query Store grids show "Showing since" and the effective start in the same case. The web viewer shows the same note on those panels. +- **`get_query_store_top` stays under the MCP response-size budget** ([#4198]): The default call returned about 48 KB, over the 32 KB limit some MCP clients enforce. `query_text` is now a 400-character preview by default. A new `full_text` argument returns the whole statement. A new `query_text_truncated` flag on each row says whether the text was cut. The web viewer is unchanged and still shows up to 2,000 characters. +- **`describe_custom_view_catalog` stays under the MCP response budget** ([#4198]): The Custom Views compose catalog returned 98 KB by default, three times the budget. It now returns a compact, source-grouped list by default (name, purpose, aggregate/unit vocabulary only). A new `source` argument drills into one source's full detail. A new `full_detail` argument returns the complete catalog as before. The web Custom Views editor is unaffected: it reads the full catalog through its own endpoint. +- **MCP `get_collection_health` caps its response size** ([#4198]): A default call on a busy SQL Server fleet reached 41,669 bytes, over the 32 KB limit. Healthy collectors with nothing to report now compact to a shorter row by default. Any failing, stale, stopped, erroring, denied, or regressed collector keeps every field. A new `full_detail` argument restores every field on request. The web dashboard and Custom Views are unaffected. +- **MCP `get_blocking` (Darling) and `get_blocked_process_reports` (Lite) cap their response size** ([#4198]): A default call measured 89,096 bytes on a 30-row test seed, over the 32 KB limit. The default row limit drops from 30 to 15. The blocked and blocking SQL text come back as a 150-character preview instead of a 2,000-character cut, and each has its own `*_truncated` flag. A new `full_text` option returns both texts whole. On Darling, a call with `dedup_key` names one incident, so it always gets the whole text. The web viewer's Blocking table and Custom Views keep 30 rows and the 2,000-character text. +- **MCP get_collection_log stays under the response-size budget by default** ([#4198]) - +- **`get_query_store_regressions` default calls stayed under the MCP response budget** ([#4198]). The +- **get_pg_io_trend's default call now fits the MCP response budget** ([#4263]): The default 24-hour call +- **`get_active_queries` stays under its response-size budget by default** ([#4261]) - A default call on a busy server returned well over the 32 KB target. The number of fields per row and an uncapped query text preview were the cause. The default page size is now 25 rows (was 50). Query text now previews to 500 characters with a `query_text_truncated` flag. A new `full_text` argument opts back into the whole text. The web viewer's Active Queries tab keeps its pre-existing 2,000-character preview. +- **get_index_usage's default answer now fits the response budget** ([#4260]) - A default call, with no `limit` argument, always returned up to 200 rows regardless of width. That measured over 69 KB on a busy server, large enough for an MCP client to refuse it outright. The default `limit` is now 75 on both Darling and Lite. An explicit `limit` still returns exactly what was asked for. +- **`get_query_heatmap` default call stays under the MCP response budget** ([#4259]): A default call returned +- **`get_object_locking` default response stays under the MCP budget** ([#4258]): Before this fix, the tool returned up to 200 locking rows with no row limit. On a busy server it pushed 71 KB, and some MCP clients refused the response. It now returns 75 rows by default, highest contention first. It reports `truncated`/`objects_returned` when there are more. Pass `limit` for more rows, up to 1000. +- **`get_plan_corrections` default call stays under the MCP response budget** ([#4198]): A default call returned nearly 100 KB on a busy server. Each row carries 20+ fields and the query text ran to 2,000 characters. The text preview is now 150 characters. Each row gets a `query_text_truncated` flag and a `full_text` opt-in. The default row limit dropped from 50 to 25. +- **Job History loads faster and stops polling every 30 seconds** ([#4256]): The tab used a window function +- **Availability Group reads: bounded the newest-snapshot lookup** ([#4228]). All five AG reads used an +- **`get_deadlock_detail` no longer returns oversized deadlock graphs by default** ([#4254]) - The tool's `deadlock_graph_xml` field could push a default call past 100 KB, because each graph runs tens of kilobytes and the tool did not cap that field's size. It now previews each graph to 2000 characters by default and marks it `deadlock_graph_xml_truncated: true`. A new `full_graph` argument returns the whole graph. A `dedup_key` call, naming one incident, always returns it in full. +- **Database Sizes and Storage Growth load without a multi-chunk planning tax** ([#4252]): The latest snapshot reads had no bound on collection time. PostgreSQL had to plan a step for every retained chunk before running each read. Each read now resolves the snapshot timestamp in a separate step first. The main read then skips chunks it does not need. Clock skew between reader and collector does not drop the latest snapshot. +- **Alert triage links now skip only a disabled dashboard, and email carries one too** ([#4230]) - the per-alert triage-page link (`web.publicBaseUrl`) used to go out even when the dashboard was off, and the web host rejected every request naming the DNS host that setting suggested using. The link is now omitted only when the dashboard is disabled. The web host accepts that configured host. Alert email now carries the same link the other channels already did, in both the HTML and plain-text bodies. +- **PLAN_REGRESSION no longer deduplicates the whole raw Query Store slice every pass** ([#4208]). On a store ingesting many servers' Query Store data, the plan-regression check was the single heaviest statement in the store. It spent most of its time sorting and spilling the server's entire raw Query Store history to temp files. This adds store schema rung V143, the `collect.query_store_interval_latest` table, holding only each interval's latest snapshot. The check now reads that table and falls back to the old full read only when coverage is incomplete for a server. Findings are unchanged: the same regressions, the same snapshots. +- **Force-plan bot no longer acts on a best plan older than 4 days** ([#4208]). The unattended force-plan bot now gates on `ForcePlanBotPolicy.MaxBestPlanAgeDays` = 4. A target whose best plan has not been seen in more than 4 days is blocked with reason `best_plan_stale`. This prevents a force on a plan the server evicted. The gate value is pinned by `ForcePlanBotPolicyTests`. +- **`audit_config` no longer runs the full analysis pass** ([#4206]). It read a full ~30-family collection, detector and scorer pass just to answer 8 point-in-time configuration facts, measured at 7-15 seconds per call. It now runs only the specific family reads those 8 facts come from. The facts and recommendations returned are unchanged. +- **`get_query_store_regressions`' comparison baseline is now a fixed 7 days** ([#4206]). The baseline used to be every Query Store capture the store retained before the requested window. Both cost and the comparison period grew with retention. It is now a fixed 7-day lookback ending at the window's start, reported as `baseline_start` and `baseline_end` in the response. A regression against something that changed more than 7 days before the window is no longer caught. A store retaining less than 7 days is unaffected. Performance Monitor Lite's own copy of this same tool had the same unbounded baseline and is fixed the same way. +- **`get_pg_cpu_utilization` now returns bucketed points instead of one row per minute** ([#4206]). A wide window (a week, or longer with the V136 host-memory columns) returned several megabytes per call. It now buckets to a point budget the same way the rest of the trend family does. It averages CPU/ACU per bucket while keeping each bucket's peak. The worst (not average) memory-pressure sample per bucket is kept, so a brief spike is not smoothed away. +- **The desktop viewer's Query Store Regressions grid had the same unbounded-baseline shape as `get_query_store_regressions` and is now bounded the same way** ([#4206]). It was a third, separate copy of the defect, in the WPF app rather than the MCP surface, with its own 7-day fixed lookback. - **The deadlock and plan-capture log patterns only offer a report whose ERROR: is the line's own label** ([#4042]) - A crafted SQL statement that echoed "ERROR: deadlock detected" plus a tab-continued DETAIL line matched the deadlock pattern through to the line's real process-id bracket. Measured end to end this forgery stored nothing, because the log assembler already takes a report's label from the line it actually came from, and new tests pin that for both the SQL and C# copies of the pattern. Both copies are narrowed anyway, matching the rule behind them exactly: the gap before ERROR: and DETAIL: can no longer cross another field's label, and the managed log-prefix family's gap to the process-id bracket can no longer slide past the real bracket into a statement's text (the same fix #4016 made for plan capture). Fewer forged lines are offered as candidates at all. - **Plan capture keeps working under a custom log prefix with fields before the process id** ([#4016]) - Plan capture required exactly one token between the timestamp and the process-id bracket in log_line_prefix. A self-hosted prefix with more than one field there (for example '%m %u@%d [%p] %Q ') matched nothing, so those stores silently stopped capturing plans. The gap before the process id now allows any number of fields, as long as none of them is itself a bracket, so a forged log line behind the real bracket still cannot be read as a plan. - **Trace flag reads no longer hide every enabled flag after an ordinary run** ([#4032]) - get_trace_flags, the web viewer's Trace Flags grid, and Lite's equivalent compared the newest captured row's timestamp against the newest successful run's timestamp, but Darling stamps a run's collection-log row when the run ENDS, after its capture rows are written - so that comparison read false after every ordinary run and hid every enabled flag. All three reads now keep the newest capture unless the newest successful run explicitly captured zero rows (which means every flag is off); a failed run, or one with no row count, keeps the reading from before this fix. @@ -387,6 +515,14 @@ Releases before 3.0.0 are not archived: those entries carry no prose to move. - **The pin that held the defect as intended behavior is renamed, which is why this was its own issue.** `DarlingMcpObjectStatsToolsTests.IndexLockingSql_PerDatabaseLatest_NonzeroWaits_ContendedFirst` asserted per-database "latest" as the read's purpose, by name; rewriting a behavior a test claims as intended is a stated design change rather than a quiet edit, and the rename is the statement. It is now `..._ServerLatestCapture_...`, and its anchor assertion pins the whole server-latest subquery instead of a bare `MAX(collection_time)` - the retired shape contained one of those too, so the old needle could not tell the fix from the defect - with the grouping asserted absent so the pin cannot pass again by going back. `ViewerFinOpsSqlTests`' three literal fragments are re-fragmented to the new SQL with the anchor pinned in all three, the database filter moving to `SqlTextPin` because the fix qualified it to `ios.database_name` when the CTE referencing the bare name went away - #3217's exact case, where an alias appearing in front of a column must not red a pin asserting a read still filters. `McpLatestSnapshotStampTests`' comment claiming this read "takes MAX per database, not one instant" was true when written and is now false; the correction is recorded rather than deleted, because the read IS one instant now, which is precisely what a `captured_at` stamp needs to be truthful and makes the A10 residual it still sits in easier to close. `StorageCommandTimeoutTests` needed no re-anchoring - its fixture pins the `GetIndexLockingAsync` ternary's C# body text and the substitution changed only the const strings that method selects between - verified by reading the method after the edit rather than assumed. - **A census caught what no pin was watching for, and its roster grew rather than shrank.** `McpPayloadContractCensusTests` rosters latest-anchored reads that do not project their anchor column, shrink-only by design, and the fix ADDED `IndexLockingSql` to it. The reason is the defect restated from another angle: the read was invisible to that census because its anchor sat inside a CTE, which the scan deliberately ignores on the grounds that such an anchor picks a SET rather than a snapshot - exactly true of per-name groups, which picked one row per name the store had ever seen. Becoming a real snapshot read is what made the census see it at all. It is rostered beside its two neighbours with that reasoning written into the roster's own comment, and all three leave together when the A10 lane stamps the family. +### Security + +- **Scrubbed legacy plan-force-action audit text** ([#4384]) - a one-time startup pass rewrites audit lines written by older builds to the same safe form the read path already applies; every other field and row is left untouched. +- **Collected PostgreSQL statement text applies the same sensitive-statement filter as the store's own statements** ([#4383]) - matching text is withheld before storage, and a one-time job rewrites rows stored earlier. +- **Stored PostgreSQL setting values are redacted in more shapes, and the one-time scrub resumes a large server-day after a restart** ([#4380]). +- **Plan-force journal reads no longer return pre-#4326 exception text** ([#4363]) - Both plan-force journal reads now replace exception text written before #4326 with a fixed sentence; evidence lines written after #4326 are returned unchanged. +- **Stopped `/api/ping` and the analysis notes from echoing raw exception text** ([#4326]) - `/api/ping`'s + ## [3.8.0] - 2026-09-17 Full entries: [docs/changelog/3.8.md](docs/changelog/3.8.md) @@ -3812,3 +3948,126 @@ Full entries: [docs/changelog/3.0.md](docs/changelog/3.0.md) [#4137]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4137 [#4141]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4141 [#4157]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4157 +[#4186]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4186 +[#4198]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/4198 +[#4203]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/4203 +[#4206]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4206 +[#4208]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4208 +[#4228]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4228 +[#4230]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4230 +[#4252]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4252 +[#4254]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4254 +[#4256]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4256 +[#4258]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4258 +[#4259]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4259 +[#4260]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4260 +[#4261]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4261 +[#4263]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4263 +[#4271]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4271 +[#4278]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4278 +[#4279]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4279 +[#4280]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4280 +[#4281]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4281 +[#4282]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4282 +[#4285]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4285 +[#4286]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4286 +[#4287]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4287 +[#4288]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4288 +[#4291]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4291 +[#4292]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4292 +[#4293]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4293 +[#4294]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4294 +[#4295]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4295 +[#4297]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4297 +[#4302]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4302 +[#4303]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4303 +[#4304]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4304 +[#4306]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4306 +[#4307]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4307 +[#4308]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4308 +[#4309]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4309 +[#4311]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4311 +[#4317]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4317 +[#4319]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4319 +[#4321]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4321 +[#4324]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4324 +[#4325]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4325 +[#4326]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4326 +[#4327]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4327 +[#4328]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4328 +[#4329]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4329 +[#4330]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4330 +[#4331]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4331 +[#4333]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4333 +[#4336]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4336 +[#4337]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4337 +[#4338]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4338 +[#4339]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4339 +[#4340]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4340 +[#4341]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4341 +[#4342]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4342 +[#4344]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4344 +[#4345]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4345 +[#4347]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4347 +[#4350]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4350 +[#4351]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4351 +[#4352]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4352 +[#4353]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4353 +[#4357]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4357 +[#4358]: https://github.com/erikdarlingdata/PerformanceMonitor/issues/4358 +[#4359]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4359 +[#4361]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4361 +[#4362]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4362 +[#4363]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4363 +[#4364]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4364 +[#4365]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4365 +[#4366]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4366 +[#4367]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4367 +[#4368]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4368 +[#4370]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4370 +[#4371]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4371 +[#4372]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4372 +[#4373]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4373 +[#4380]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4380 +[#4382]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4382 +[#4383]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4383 +[#4384]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4384 +[#4386]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4386 +[#4387]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4387 +[#4388]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4388 +[#4390]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4390 +[#4391]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4391 +[#4392]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4392 +[#4393]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4393 +[#4395]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4395 +[#4396]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4396 +[#4399]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4399 +[#4400]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4400 +[#4401]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4401 +[#4402]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4402 +[#4405]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4405 +[#4406]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4406 +[#4407]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4407 +[#4408]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4408 +[#4409]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4409 +[#4411]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4411 +[#4412]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4412 +[#4413]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4413 +[#4416]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4416 +[#4417]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4417 +[#4418]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4418 +[#4419]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4419 +[#4420]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4420 +[#4421]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4421 +[#4422]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4422 +[#4423]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4423 +[#4424]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4424 +[#4429]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4429 +[#4430]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4430 +[#4431]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4431 +[#4432]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4432 +[#4433]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4433 +[#4434]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4434 +[#4435]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4435 +[#4436]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4436 +[#4437]: https://github.com/erikdarlingdata/PerformanceMonitor/pull/4437